17  Evaluating the works

Blue Mountains City Council Healthy Waterways — statistical analysis

17.1 What this chapter is for

The question you actually asked is whether the stormwater treatment works have improved the creeks. This chapter answers it as far as the archive allows, which is not far, and says precisely what is missing.

Three things come out of it. Dates. No treatment asset in either database carries a construction date anyone can rely on, and one project report is the only source of real ones. Design. Even with perfect dates, one catchment cannot answer this question at any level of effort — a point worth settling before the next evaluation is commissioned rather than after. A null that carries no information, which is a different thing from a null that says the works did nothing.

Every intervention date below is read at render time from a single file. Those dates are provisional, they are the weakest part of the analysis, and replacing them is a one-file job — Section 17.7 sets out how.

17.2 The data behind this chapter

Each question and request below is set out again in What we need from you, with what it blocks, what an answer is worth and what it would cost you to find, ranked against every other ask in the report.

17.2.1 Questions only you can answer

What do these fields actually record — build, handover, capitalisation, load date or revaluation batch? Construction Date in Civica, AssetPit.Year, tblAIMOpenChannel.Datecreate and Effectdate, and Asset_Audit.OC_CommDate.

OC_CommDate is the largest populated date field we have — 204 of 254 treatment assets — and whether it is usable or junk turns entirely on this. Our reading is that it is a year-end capitalisation stamp: it has only eleven distinct values, every one falls on 29 or 30 June, 169 of them are 2016-06-29, and it disagrees with all five independently documented construction dates. But that is inference, and somebody in Assets can just tell us.

Refer to it as dq:civica-date-semantics.

Given that a four-site comparison cannot detect anything, which do you want — evaluate the works on a response that varies less between sites, pool several works programs to reach the site count, or accept in advance that the evidence will be descriptive and say so?

Detecting half a score point needs roughly twelve treated and twelve control sites, sampled twice a year, for six years either side. At four per arm it is unreachable at any amount of sampling — the floor is 0.73 score points. The Leura Falls difference-in-differences returned +0.11 with a confidence interval from -0.36 to +0.58: a null that carries no information at all. Deciding this in advance is what stops a four-site study being run and then reported as “no effect”.

Refer to it as dq:stormwater-eval-design-decision.

Would it be possible to flag a sample in the database when something has happened to the creek — a spill, a contamination, a fire, works in the channel — rather than leaving it to a free-text comment or to nothing at all?

The July 2012 bifenthrin contamination of Jamison Creek destroyed the macroinvertebrate community for three months and is completely invisible in the database. Twenty-four samples across the impact and recovery period carry no marker of any kind, and anyone analysing the Jamison series without knowing the history would read a catastrophic contamination as a poor run of monitoring results. The 2023 Hazelbrook incident is recorded, but only as two sentences in a site note. This one recovers nothing already lost - it is about the next incident, not this one.

Refer to it as dq:flag-incidents-against-samples.

17.2.2 What would answer them

Where did the Leura Falls monitoring data end up? Sites E151 and E195-E199, six sites monthly from June 2014 to July 2017, plus rain-event inlet and outlet sampling of five treatment systems with measured pollutant reductions per system.

Not one E-code site exists in either database. This is the only dataset anywhere that could say whether an individual treatment device works — which is a different question from the network-level one the commissioning dates would unlock, and arguably the more useful of the two. It is your own data, recently collected, and somebody knows where it is.

Refer to it as dq:leura-falls-monitoring.

When was each treatment asset built? The specific ask: a OneCouncil / CiAnywhere extract filtered to SWDR-GPTS and pit types 07, 08, 11 and 12, returning the FA450 financial-asset commissioning date — plus the DA and construction-certificate records, which have never been searched, and whoever populated “EConst. Date” on 28 SQIDs.

Your central question is whether the stormwater works have improved waterway health, and it cannot be answered because nothing reliably records when the works happened. This would take the evaluation from three usable comparison units to something near sixty, and the smallest change it could detect from more than half a score point to about a fifth of one. Be clear about what it does not do, though: dates make the question askable, not answerable. The monitoring design decides the rest, which is the next item but one.

Refer to it as dq:sqid-commissioning-dates.

Which catchment — or at minimum which upstream drainage node — does each treatment asset serve? It exists nowhere except for the seven Leura Falls systems.

Without it, “downstream of a treatment device” cannot be defined at all, so no monitoring site can be paired to any asset and no asset can be evaluated regardless of how good its date is. This is the quieter half of the commissioning-date problem and it is just as binding.

Refer to it as dq:sqid-served-catchment.

Do you still hold the two datasets that sit beside the 2012 Jamison Creek macroinvertebrate sampling — the OEH bifenthrin laboratory results (8 sites, 19 samples, July 2012) and Robert McCormack’s freshwater crayfish surveys (3 sites, 30 samples, July 2012 to April 2014)?

Your conference paper on the incident reports all three datasets together, but only the macroinvertebrate sampling reached the database we have. The bifenthrin results would give us an actual dose against an actual response, which nothing else in twenty-six years of macroinvertebrate monitoring can offer, and the crayfish surveys track the species the kill was noticed by. Both would turn a well-documented incident into a quantitative exposure-response series on one of your own creeks.

Refer to it as dq:jamison-2012-companion-data.

Is there a list of the treatment assets that are not pits — rock-lined biofilters, raingardens, bioretention basins, sediment basins, trash racks — and GIS layers for them? They are largely invisible to the asset system.

We have a partial list already: sqid_dates.csv carries 315 treatment assets, including biofilters and rock-lined channels picked up from your 2021 Stormwater System Maintenance Program spreadsheet that are not in the GIS pit table. What is missing is a register that says the list is complete, so we cannot tell whether the gap is five assets or a hundred.

Refer to it as dq:non-pit-assets.

Where would a decommissioning or major-maintenance date be recorded — is there an asset-disposal workflow in Civica/OneCouncil at all? The Disposed extract returns zero SQIDs, which we read as “none are captured” rather than “none have ever been removed”.

Same mechanism as the commissioning dates and smaller: without them we cannot tell whether an asset was actually operating during the monitoring period it is being credited with.

Refer to it as dq:sqid-decommissioning.

Does a fine-grained stormwater sub-catchment layer exist, or an imperviousness layer to go with it? Data\Environment on the corporate GIS share is ACL-denied to us and its orphaned label layers point straight at sub-catchment polygons — could we be given read access to it?

It is what would let the drainage area served by an asset be defined at anything finer than the delineated site catchment, which is currently the only unit available and is far too coarse for a device serving a few streets.

Refer to it as dq:stormwater-subcatchment-layer.

Would you ask Sydney Water for records dating the original Vale St wetland cells? They were built by Sydney Water in the 1990s rather than by Council, which is why they are not in your register.

One asset, but a load-bearing one: it is the intervention date for the one catchment with a defensible before-and-after case.

Refer to it as dq:sydneywater-valest.

17.3 What the record holds

The methods document lists stormwater treatment projects in the Leura Falls, Jamison, Kedumba and Knapsack Creek catchments as one of the reasons sites were selected. The databases record almost nothing about them.

Table 17.1: Every provenance note in the site register that bears on a project or an incident. These few strings are the entire record of why sites were added.
Site Waterway Note
09.2BBH Centennial Glen Creek New site for Centennial Glen - shifted 2024 due to safety issues & poor habitat at 09BBH
59.2BLA Leura Falls Trib u/s Chelmsford Dr replacing 59BLA (Cumberland), due to private land access issues
84NBX Cripple Creek Added 2021
85BKT Kedumba Creek Added 2022 for Kedumba Project
86BKT Kedumba Creek Added 2022 for Kedumba Project
87GHZ Hazelbrook Creek tributary New site added as control for Hazelbrook bifenthrin contamination incident 2023
88GHZ Hazelbrook Creek tributary New site added 2023 - impact site for bifenthrin contamination
P2-M5 Bedford Tributary Old Site not sampled after 2005

tblProjects in the water quality database lists Macro, RecWQ, Sediment, Jamison and Leura as projects, but only Macro and RecWQ are ever used: the Jamison and Leura catchment restoration projects have zero samples tagged. Project attribution therefore has to be inferred from site codes and dates, and no commissioning date for any treatment device exists in either database, in the methods document, or in Council’s public material.

There is one designed impact study in this archive, and nothing in the databases says so. In July 2012 more than a thousand dead Giant Spiny Crayfish turned up along two kilometres of Jamison Creek at Wentworth Falls. Bifenthrin had been over-applied at a residential development on 5 July, about 3 mm of rain fell on the 6th, and the kill was found the next day; the pesticide reached the creek 300 m away through a stormwater pit and pipe run, and two pest control operators were later convicted. All of this is in your own conference paper (St Lawrence et al. 2014) — none of it is in the monitoring database.

The sampling is there, though, and it is the closest thing in the archive to a designed impact study. Monthly sampling ran through the impact period at three sites: 61BWF 290 m upstream of where the pesticide entered, which is the control; 62BWF 120 m downstream and 23.2BWF 2.2 km downstream, which are the impact sites. The dates match the paper’s figure exactly. From 2013 the three drop back to the network’s routine cadence of twice a year, and the record does not stop there: the last visit is 3 April 2024. The signal is not subtle — across July, August and September 2012 the control holds 7, 13 and 16 families while the near site holds 1, 3 and 3.

It is not, however, a before–after–control–impact design. 61BWF and 62BWF were both first visited on 12 July 2012, five days after the kill was found: neither the control nor the near impact site has a single macroinvertebrate sample from before the over-application, and the water quality database — a separate record that could in principle have supplied one — holds none at either site either. The only pre-incident sampling anywhere in the set is at 23.2BWF, 2.2 km downstream: five samples between 3 April 2008 and 14 March 2012, the last of them four months before the spill. The Before × Control cell is therefore empty, and the interaction that defines a BACI — how much more the impact sites changed than the control did — cannot be estimated at all. What the archive actually holds is an after-only control–impact design with nineteen rounds of paired temporal replication, plus a before-and-after series at one of the two impact sites.

That is a limitation of the event, not of Council’s response. A pesticide spill is not scheduled, and both monitored sites exist because of it; no amount of diligence produces a pre-incident sample at a site established afterwards. The upstream reach is a spatial control — the standard and correct design for an impacts-and-recovery study of an unannounced contamination, which is exactly what St Lawrence et al. (2014) set out to do. The design is right for its purpose. It is simply not the design that would let this report difference out a catchment trend.

The control also drops out before the impact sites do, and anyone treating this as a control–impact comparison has to stop where it does. 61BWF was last sampled on 19 April 2021; 62BWF and 23.2BWF were sampled six more times after that, out to 3 April 2024. Those last six rounds have no contemporaneous upstream comparison, so they can only be read against the impact sites’ own history. The paired part of the design ends in 2021, and a recovery trajectory fitted past that year is an impact-site time series with the control quietly dropped, which is the inference the next section explains a control exists to prevent.

That matters here for two reasons beyond the incident itself. It is the only place in the 26 years of macroinvertebrate monitoring — 1998 to 2024, which is not the same span as the 27-year water quality record — where we can see what a known dose of a known contaminant does to these creeks and how long recovery takes. And the mechanism is the one this whole report keeps circling: the pesticide travelled from a property to a creek 300 m away through the stormwater network, which is what a connected, impervious catchment is. The paper’s own conclusion is about effective imperviousness, and it is the only published test of that mechanism on a Blue Mountains creek.

17.3.1 Could a before–after–control–impact design work at all?

A BACI analysis needs four things: samples at the treated site before the works, samples after, matched control sites monitored over the same period, and a known intervention date. The requirement is not a formality invented here — without a control, an impact cannot be separated from whatever else happened between the two periods, and without repeated sampling in each period the comparison has no error term to test against (Stewart-Oaten et al. 1986). Underwood (1994) goes further: a single control is not enough either, because a treated and a control site can drift apart for reasons unrelated to the works. That drift is exactly the variance component that dominates the power calculation in Section 17.5, so it is named here rather than there.

Table 17.2: Macroinvertebrate sampling coverage at sites in the Jamison, Kedumba and Leura Falls catchments. Knapsack Creek, the fourth catchment the methods document names, carries no treatment asset at all in the stormwater extract and is not part of this design; the note below says why.
Catchment Site Waterway First year Last year Years sampled Samples
Jamison 23BWF Jamison Creek 1998 2007 7 15
Jamison 22BWF Valley of the Waters 2000 2013 10 19
Jamison 24BWF Wentworth Falls Lake 2006 2024 18 33
Jamison 23.2BWF Jamison Creek 2008 2024 16 31
Jamison 61BWF Jamison Creek 2012 2021 10 20
Jamison 62BWF Jamison Creek 2012 2024 13 25
Jamison 22.2BWF Lillians Glen 2015 2024 10 10
Kedumba 19BKT Kedumba Creek 1998 2024 23 38
Kedumba 16GKT Katoomba Creek 2000 2024 21 28
Kedumba P2-U8 Katoomba Creek 2000 2004 2 4
Kedumba 73BKT Kedumba River 2016 2023 5 6
Kedumba 85BKT Kedumba Creek 2022 2024 3 5
Kedumba 86BKT Kedumba Creek 2022 2024 3 5
Leura Falls 21BLA Gordon Creek 1998 2024 20 29
Leura Falls 20BLA Leura Falls Creek 1999 2024 22 46
Leura Falls 58BLA Leura Falls Creek 2011 2024 13 22
Leura Falls 59BLA Leura Falls Creek 2014 2021 8 16
Leura Falls 59.2BLA Leura Falls Trib u/s Chelmsford Dr 2023 2024 2 3

Three of the four named catchments, and the fourth is accounted for rather than dropped. Section 17.3 names Leura Falls, Jamison, Kedumba and Knapsack Creek; Table 17.2 carries the first three. Knapsack Creek is monitored, and monitored well — 27 samples at 1 site (50NEP) across 19 years from 2002 to 2024, a longer record than 12 of the 18 series in Table 17.2. It is left out because there is nothing in it to evaluate: the stormwater extract places no treatment asset at all in that site’s catchment group, while the only asset in the whole 315-asset extract whose location names Knapsack is a gross pollutant trap carrying nothing but a 1997 drainage-inventory date, which pre-dates the monitoring record and is excluded on the same rule as the rest of that block (Section 17.4). Read that as a hole in the asset record rather than as evidence that nothing was built there — the methods document says projects exist in the catchment, and the asset extract does not show them, which is one more thing dq:sqid-commissioning-dates would settle.

Kedumba: no. Sites 85BKT and 86BKT were, on your own note, “Added 2022 for Kedumba Project” (Table 17.1). They have 10 samples between them, all from 2022 onward. There is no before period, and no statistical method creates one. The only Kedumba sites with a long record (19BKT at Katoomba Cascades, 16GKT on Katoomba Creek) sit downstream of a catchment that has had continuous incremental works, so they integrate everything and isolate nothing.

Leura Falls: not as designed, but closest to feasible. 20BLA at Leura Cascades has 46 samples across 22 years from 1999 to 2024 — a genuine long series spanning any plausible works period. What is missing is a date and a control.

Jamison: confounded by a relocation and by the 2012 contamination. The long Jamison Creek series runs 23BWF (1998–2007) then 23.2BWF (2008–2024), a relocated pair that must not be spliced without comment — these are the paper’s ‘Armstrong St’ and ‘Weeping Rock’, 560 m and 2.2 km below the point where the bifenthrin entered. The join sits inside any plausible works window, so a step at the splice cannot be told apart from a step from the works, and the July 2012 contamination sits inside it too and would swamp either.

The design as recorded therefore cannot demonstrate whether the works changed waterway health — not because the sites are poor, several are excellent long series, but because the two things that turn a monitoring series into an experiment are absent: when each device was commissioned, and which sites are untreated controls. Two partial fixes have since arrived: the stormwater asset extract supplies dates of a sort (Section 17.4), and your own 2018 project report supplies real construction windows for seven systems in one catchment (Section 17.6). A narrow evaluation is now possible. It returns nothing, and the reason it returns nothing is that it was never capable of returning anything smaller than an improvement larger than any plausible treatment effect.

17.4 What the asset register can and cannot date

Your stormwater extract carries 315 treatment assets with a date of some kind against them. Almost none of those dates means what an evaluation needs it to mean.

Table 17.3: What survives of the stormwater asset date register. The rows are not a single funnel: the 1997 inventory and the coordinate count cut across the date classes above them. A “built” date is one whose recorded basis is construction, practical completion or a construction year; an inspection date says only that the asset existed by the day someone looked at it, which is an upper bound and nothing more. Only 24 of 315 assets carry a construction date anyone recorded as high confidence.
Stage Assets
Treatment assets with any date at all 315
… whose date plausibly means “built” 77
… of those, high confidence 24
… whose date is an inspection (upper bound only) 227
… with no usable date 11
In the 1997 drainage inventory (pre-date the record entirely) 74
With a coordinate, so placeable in a catchment 265

Three features of Table 17.3 decide what can be attempted.

Most dates are inspections. 227 of the 315 assets are dated by the first or only time somebody inspected them, which bounds construction from above and nothing more. Treating one as an intervention date would place the works at whatever moment record-keeping happened to start — for many assets 2021 or 2022, years after they were built. Three of the seven systems below carry 2021 inspection dates against construction your own project report puts in 2016.

A large block pre-dates the record entirely. 74 pit-based assets carry a 1997 drainage-inventory date. The macroinvertebrate record begins after that, so they have no before period and are excluded throughout. This is also what removes Knapsack Creek from the BACI design. The extract places 0 treatment assets in the catchment group of 50NEP, and the only asset in the whole extract whose location string names Knapsack — a gross pollutant trap by Knapsack Park at Glenbrook — is in this 1997 block, dated by an inspection and flagged low confidence. So the fourth catchment the methods document names has nothing datable in it to evaluate.

The accounting dates are not build dates. The Civica extract offers a commissioning date for all 204 treatment assets in the asset audit, and every one of the 204 falls on 29 or 30 June — financial year end — with 169 in 2016, the year of an asset revaluation. They are capitalisation dates: the Glenbrook Lagoon gross pollutant traps are stamped 1 January 2015 against notification forms showing construction in April and May 2014. Nothing here uses them.

17.4.1 The size of the prize

Table 17.4: The whole network, not one catchment. Assets counts distinct treatment assets; site pairings counts asset-and-catchment-group pairs, and is the larger number because catchment groups nest, so one device upstream of several monitoring sites is one asset and several pairings. The asset is what somebody has to find a date for; the pairing is the comparison unit an evaluation would get. Every catchment group here contains a monitoring site by construction, and the group column does not partition — a group holding assets of two date classes is counted in both rows. These are not the 315 of Table 17.3, which counts assets carrying a date of any kind wherever they sit: only 265 of those have a coordinate at all, and fewer still fall in a monitored catchment. The bottom three rows partition the 169 candidate assets by what their recorded date actually is, and the bottom two are the ones a single data request would move.
Stage Assets Site pairings Catchment groups
Treatment assets placed in a monitored catchment 214 509 62
… that do not pre-date the monitoring record 169 413 59
… of those, dated to a construction event 49 118 27
… dated only by an inspection (upper bound) 109 268 49
… carrying no date at all 11 27 15

Table 17.4 is the most important thing in this chapter. 120 of the 169 candidate treatment assets — 71% — are blocked purely by date quality. Not by geography, not by monitoring coverage, not by the works being too small: by nobody having written down when they were built. Counted as site pairings instead — the unit an evaluation actually works in, since one asset upstream of several monitoring sites offers several comparisons — it is 295 of 413, or 71%: the same share either way, so nothing here turns on which unit you count in. Those assets sit in 59 catchment groups, against the 3 treated sites the Leura Falls evaluation actually has, and there are a further 58 catchment groups in the network with no treatment asset at all to serve as controls.

Installation dates are what the evaluation question Council asked for is waiting on (dq:sqid-commissioning-dates), and Section 17.5 says what they would buy, in the units the design is measured in. They are not the report’s single highest-value request: chapter 2’s master list rates 14 asks transformative and ranks this one 12 of 107 once effort is set against value. This chapter’s own dq:leura-falls-monitoring ranks above it, and it should — dates make the network question askable, while the Leura Falls data is the only thing in existence that could evaluate a single device.

17.5 How much monitoring an evaluation needs

Table 17.5: Variance components of the average factor score, as standard deviations in score points, fitted on edge samples in six-year periods (257 samples at 13 sites in the Jamison, Kedumba and Leura Falls treatment catchments, Knapsack Creek excepted as above; 1298 samples at 77 sites network-wide). Only the last 3 rows enter a difference-in-differences power calculation, and they enter it differently: tau is divided by the number of sites, sigma_year by the number of years, and sigma by the number of samples.
Component Treatment catchments Whole network
Between sites 0.57 0.68
Shared between periods (differenced out by a control arm) 0.22 0.38
Site x period — how differently sites move (tau) 0.18 0.10
Site x year — how much one site moves year to year (sigma_year) 0.29 0.33
Within site and year — sampling noise (sigma) 0.47 0.45
Table 17.6: Power of a difference-in-differences evaluation of a stormwater treatment works, at 5% significance, with equal numbers of treated and control sites each sampled twice a year for six years either side. Power is simulated with the simr package (500 simulations per row) and tested on Satterthwaite degrees of freedom; the Smallest detectable change column is the closed-form equivalent of that simulation on the same reference distribution, shown to confirm the two agree. One rating class is 1.00 score points wide. The Floor (unlimited sampling) column is the same closed form with the sampling term removed, so it is a limit rather than a check: nothing in the simulation confirms it, and it does not fall to zero, because two of its three components are not sampling noise — sites differ in how much they change between two periods (tau = 0.18), which is divided by the number of sites, and one site moves from year to year (sigma_year = 0.29), which is divided by the number of years. Only sigma = 0.47 is divided by the number of samples.
Sites per arm Power to detect 0.50 95% CI Smallest detectable change Floor (unlimited sampling)
4 40% 35-44% 0.86 0.73
8 75% 71-79% 0.54 0.46
10 79% 76-83% 0.48 0.41
12 91% 89-94% 0.43 0.37
14 96% 93-97% 0.40 0.34
16 98% 97-99% 0.37 0.31

Two sources of variation matter here, and counting only one of them is the commonest way to under-resource a monitoring program.

The first is sampling noise: a single sample at a single site is a noisy measure of that site’s condition, standard deviation 0.47 units (Table 17.5). More sampling fixes that. The second is that sites do not all change by the same amount between two periods, even without any treatment — a creek recovers from a fire, another silts up, a third changes for no reason anyone recorded. Measured on these very sites the standard deviation of that movement is 0.18 units, and the whole-network estimate of 0.10 says it is not a quirk of the three treatment catchments. Because a difference-in-differences contrast is a difference of site-level changes, this second component is divided by the number of sites, not the number of samples. No amount of extra sampling reduces it.

There is a third, and leaving it out is what makes a power calculation flatter itself. A single site also moves from one year to the next — standard deviation 0.29 units here, 0.33 network-wide — and two samples taken in the same year share that movement. Counting them as two independent draws is the arithmetic that turns twelve samples over six years into twelve independent measurements when they are nearer to six. That component is divided by the number of years, so it is reduced by a longer record and not by a busier one.

What that costs:

  • Four treated and four control sites, sampled twice a year for six years either side, has only a 40% chance of detecting an improvement of half a score point that genuinely occurred. The change it could reliably detect is about 0.86 score points.
  • At four sites per arm, half a score point is out of reach at any amount of sampling: the floor is 0.73 score points.
  • Detecting half a score point needs about 12 treated and 12 control sites, each sampled twice a year for about six years either side. That reaches 91% power, and the closed form agrees at 0.43 score points.
  • A change of a quarter of a score point is not worth designing for: it needs at least 25 sites per arm on the floor alone.

Judgement call 7 is answered by Table 17.6, and the answer is unwelcome: one catchment cannot answer this question at any level of effort. If 12 sites per arm is beyond reach — and for most councils it will be — the conclusion is not to run a smaller version of the same design. It is that a difference-in-differences on the health score is the wrong instrument, and you would do better to evaluate the works on a response that varies less between sites (a targeted water quality or physical habitat measure at the outfall), or to accept in advance that the evidence will be descriptive rather than statistical. Designing a four-site study and reporting its null as “no effect” is the worst of the three.

17.5.1 What fixing the dates would buy

Table 17.7: What installation dates are worth, in score points. The comparison unit is a monitoring site in the top row, which is the design as it actually stands, and a catchment group with a monitoring site in the two rows below it; in neither is it a device. Figures are the closed-form minimum detectable effect at 80% power and 5% significance on a t reference, using tau = 0.18, sigma_year = 0.29 and sigma = 0.47 as measured over six-year blocks in the treatment catchments. One rating class is 1.00 score points wide. The middle row is what could be attempted today; the bottom row is the ceiling a complete date register would put on it.
Scenario Treated units Control units Smallest detectable, 12 samples each side Smallest detectable, 6 samples each side Floor (unlimited sampling)
Now — the seven Leura Falls systems 3 9 0.75 0.85 0.64
Catchments with an asset already dated to a construction event 27 58 0.24 0.27 0.20
Every candidate asset dated 59 58 0.19 0.21 0.16

So dating the assets moves the evaluation from a design that cannot see 0.65 score points to one that could find about 0.19 score points, with a floor of 0.16 that no extra sampling removes (Table 17.7). That is about a fifth of a score point, and it is the region where a real catchment-scale treatment effect might plausibly live.

Two honest riders, because this is the number the whole data request rests on.

It is a ceiling, not a forecast. It assumes every one of the 59 catchments is datable, that each has twelve edge samples either side of its own intervention, and that the intervention dates are far enough apart not to collide with the before periods of the others. None of those is guaranteed; the real figure will be worse.

The brief’s range describes the floor, not the evaluation. The floor of 0.16 does sit inside the “0.1 to 0.2” the brief claimed, so on magnitude the brief is close. What it gets wrong is what the number is. 0.16 is an asymptote: it is set by tau alone, it is what this design converges on as sampling goes to infinity, and there is no budget, on this response and with this many catchments, that betters it. The figure a real program would achieve is the one above it — 0.19 at twelve samples a side — and the rider above makes even that a ceiling rather than a forecast. So quote 0.19, name 0.16 as the floor beneath it, and do not quote a range beginning with 0.1: nothing in this design reaches 0.1.

17.6 The seven Leura Falls systems

Between September 2015 and June 2017 seven stormwater treatment systems were built in the Leura Falls Creek catchment under the Leura Falls Catchment Improvement Project (Table 17.8). They are the only treatment assets anywhere in the network with a recorded served catchment, and the project’s own final report (Blue Mountains City Council 2018) states a construction window for each one. That report is the source of every date used here.

Table 17.8: The seven Leura Falls systems, as dated by the 2018 project report. Constructed is the end of the construction window the report gives, because a system treats nothing until it is finished. The report writes “constructed around …” for every one of the seven, so none of these dates is better than month precision. The Vale St catchment figure covers both flowlines. These dates, and only these, are what the evaluation below runs on; they live in R/data-layer/stormwater_intervention_dates.csv and can be replaced without touching any code.
System Town Catchment (ha) Impervious Constructed Confidence
Vale St southern flowline Katoomba 24 55% Apr 2016 medium
Vale St northern flowline Katoomba Jun 2017 medium
Jersey Ave Leura 4 60% Mar 2016 medium
Craigend St Leura 1 60% Mar 2016 medium
Murray St Leura 5 70% Mar 2016 medium
Kanimbla St Katoomba 11 30% Sep 2015 medium
Commonwealth St Leura 6 50% Oct 2016 medium

17.6.1 Which monitoring site is downstream of which system

A treatment system only counts as an intervention at a monitoring site if the ground it treats actually drains to that site. That is a drainage question, not a distance question, and it is settled here against the DEM-delineated catchments rather than by proximity. Six of the seven systems have no coordinate in any asset register — the rock-lined biofilters and raingardens are largely invisible to the asset system — so each system is represented by its road corridor, taken from the AssetRoadSurface layer, and the evidence published for each pair is how much of that corridor falls inside the site’s catchment.

Table 17.9: Every system-to-site pair the drainage evidence supports, at a threshold of 10% of the road corridor inside the catchment. This is the assumption the whole evaluation rests on and it is weaker than a traced pipe connection: it establishes that runoff from the system’s street reaches the site, not that the treated flow does. A reviewer wanting to check one pair should start with the corridor fraction.
System Monitoring site Site catchment (ha) Corridor in catchment
Commonwealth St 20BLA 215 137 m of 137 m (100%)
Craigend St 20BLA 215 239 m of 2058 m (12%)
Craigend St 21BLA 175 1190 m of 2058 m (58%)
Jersey Ave 20BLA 215 586 m of 790 m (74%)
Kanimbla St 20BLA 215 706 m of 706 m (100%)
Kanimbla St 58BLA 73 577 m of 706 m (82%)
Murray St 20BLA 215 344 m of 344 m (100%)
Vale St northern flowline 20BLA 215 395 m of 556 m (71%)
Vale St northern flowline 59.2BLA 66 395 m of 556 m (71%)
Vale St northern flowline 59BLA 39 395 m of 556 m (71%)
Vale St southern flowline 20BLA 215 395 m of 556 m (71%)
Vale St southern flowline 59.2BLA 66 395 m of 556 m (71%)
Vale St southern flowline 59BLA 39 395 m of 556 m (71%)

Two consequences fall straight out of Table 17.9, and both cost the evaluation something.

The obvious control is not a control. Gordon Creek (21BLA) is the site you would otherwise reach for: a long series in the same town, outside the Leura Falls Creek catchment. But 58% of the Craigend St road corridor drains to Gordon Creek, so Gordon Creek is downstream of a treatment system and belongs in the impact arm, not the control arm. It is a poor member of either: Craigend St treats 1.0 ha of a 175 ha catchment, or 0.6% of it (Table 17.10). It is dropped from both arms.

The Commonwealth St system is downstream of the Commonwealth St site. Site 58BLA is named for Commonwealth St, but the Commonwealth St road corridor lies outside 58BLA’s catchment. 58BLA is upstream of that system and cannot be used to evaluate it. It is downstream of Kanimbla St, and is used for that.

Table 17.10: How much of each site’s catchment the systems upstream of it actually treat. This is the closest thing to a dose the archive supports. A site whose catchment is 0.6% treated cannot be expected to show anything, and a site whose catchment is 61% treated is the one to watch. Systems upstream is a count of systems and matches this site’s rows in Table 17.9 one for one. Served catchments is what the treated hectares are summed over, and it is smaller wherever the two Vale St flowlines both reach a site: the project report records one served catchment for the pair, against the southern flowline, so counting it twice would double the 24 ha it serves.
Site Catchment (ha) Systems upstream Served catchments Treated (ha) Share of catchment treated
59BLA 39 2 1 24 60.8%
59.2BLA 66 2 1 24 36.4%
20BLA 215 7 6 51 23.7%
58BLA 73 1 1 11 15.0%
21BLA 175 1 1 1 0.6%

17.6.2 The result

Table 17.11: Edge-sample coverage at every site downstream of a Leura Falls system, split at the construction window (September 2015 to June 2017). Samples inside the window are discarded rather than assigned to a period, because the systems went in one at a time across it.
Site Before During After Verdict
59BLA 4 3 9 usable
59.2BLA 0 0 3 excluded — too few samples before
20BLA 20 3 14 usable
58BLA 6 4 12 usable
21BLA 13 2 7 excluded — 0.6% of catchment treated

Of the 5 sites with a treatment system upstream, 3 can carry a before-and-after comparison at all: 59BLA, 20BLA, 58BLA (Table 17.11). 59.2BLA, the relocated replacement for 59BLA, has no samples before the works and can never contribute — the same defect as the Kedumba sites, for the same reason. Three impact sites, against the 12 per arm Section 17.5 says an evaluation that could detect half a score point needs.

The difference-in-differences estimate of the effect of the Leura Falls systems on the health score is +0.11 score points (95% CI -0.36 to +0.58), from 213 edge samples at 3 treated and 9 control sites, fitted with site, year and site-by-period random effects. One rating class is 1.00 score points wide. That is a null result, and the interval is wide enough to contain both a deterioration of 0.36 score points and an improvement of 0.58.

The detectable effect size is the finding. At 80% power and 5% significance, this design could not have detected an improvement smaller than 0.65 score points — 65% of a score point, a jump most of the way from one rating class to the next. No stormwater treatment program anywhere would be expected to deliver that at the catchment scale. The observed estimate is 17% of it. The same closed form as Table 17.7, evaluated at this design’s own record rather than at the round six and twelve tabulated there — 8.9 samples per site and period, over 6.6 years — is more pessimistic again — 0.78 score points, of which 0.63 is floor — and both are more than half a score point, which is the only comparison that matters.

So the honest reading of the null is not “the works did not improve the creek” but “this design could not have told you either way” — and Section 17.5 said so in advance, before any of these data were fitted.

17.6.3 What the null is and is not confounded with

Table 17.12: Dissolved oxygen samples at the sites used here, by period and probe. The before period is entirely Hydrolab and the after period contains none, so on a probe-measured parameter a plain before-and-after comparison at these sites is a comparison of two instruments.
Period hydrolab aquaread aquatroll
before 71 0 0
after 0 24 75

The intervention window sits directly on the first of the two dissolved-oxygen instrument steps Section 3.7.2 identifies. Period and instrument era are the same variable here (Table 17.12), and no model can separate them from a single arm. This is exactly why the design needs controls rather than a before-and-after: the control sites went through the same probe replacements on the same dates, so the period-by-role interaction differences the instrument out even though the period main effect is uninterpretable. It is also why the era term cannot be carried as a covariate here the way Section 6.4 carries it — there is nothing left for it to estimate.

The health score, which carries the primary result, is not exposed at all: it is built from families identified under a microscope, not from a probe. Dissolved oxygen, pH, electrical conductivity, temperature and turbidity are fully exposed and are only interpretable through the control arm. Alkalinity, phosphate, nitrate and faecal coliforms come from test kits and are untouched by the probe changes. One of those four carries a different unknown, and it is about the analyte rather than the instrument: nothing in the record says whether phosphate is reported as PO4 or as P, and the two differ by a factor of three (Section 7.5). That is not a further confound here — it is one constant factor on one column, so it cancels out of a before-and-after contrast taken within it — but it is why nitrate-N declares its species in its name and phosphate does not.

Two further confounds land inside the after period, and neither can be differenced away by a control site outside the catchment, because both are specific to Leura Falls Creek:

  • A raw sewage leak ran into Leura Falls Creek for at least 22 months, from April 2016 to February 2018 (Blue Mountains City Council 2018). It was reported to Sydney Water and the EPA in April 2016; the source was not located until your own officers found it in February 2018. It entered the creek downstream of the upper monitoring sites but upstream of Leura Cascades (20BLA) and the Commonwealth St site, two of the three impact sites. The leak begins one month after four of the seven systems were finished and runs through the whole early after period.
  • A landscaping supplies business was discharging enough phosphorus and suspended solids into the upper catchment for a Prevention Notice to issue in September 2017 (Blue Mountains City Council 2018).

Both push the after period in the opposite direction to the treatment. A null in the presence of a 22-month sewage leak is not evidence that the works did nothing; it is evidence that the design cannot see past a much larger signal. The 2020-onward wet years (Section 4.4) overlap the after window too, though that one the control arm does difference out.

Your own 2018 report reached the same conclusion from the same catchment. It records “no clear improvement in sub-catchment-wide downstream water quality after construction of stormwater treatments”, and “no clear evidence of improvement in waterway health ratings” (Blue Mountains City Council 2018). What is added here is a control arm and a stated detectable effect size, which is the part that turns a disappointment into a design specification.

Table 17.13: Difference-in-differences on the probe-measured water quality parameters, with the instrument change differenced out by the control arm. Every interval spans zero and every detectable effect size is large. For scale, the flow-state effect chapter 6 establishes on dissolved oxygen (Table 6.15) is 2.70 points of saturation — a real and well-estimated effect more than an order of magnitude smaller than what this design could have found.
Parameter n Treated sites Effect 95% CI Smallest detectable
Dissolved oxygen (% saturation) 152 2 -11.48 -40.37 to +17.42 39.90
Dissolved oxygen (mg/L) 159 2 -2.00 -5.20 to +1.20 4.42
pH 164 2 +0.11 -0.49 to +0.72 0.84

Only 2 treated sites survive the water quality coverage requirement (Table 17.13), against 3 for the health score, so these are weaker still.

17.7 If better dates arrive

Every intervention date in this chapter is read at render time from stormwater_intervention_dates.csv, in R/data-layer/. Nothing downstream of that file contains a date; the cached spatial matching holds none either, so it does not need rebuilding when one changes. Replacing a date and re-rendering moves every number, table and sentence above that depends on it.

That is deliberate, because the dates are provisional and should be read that way. The seven systems are dated from a report that writes “constructed around” in front of each one, and the rest of the register is inspections and capitalisation dates. If Healthy Waterways or Assets hold commissioning records, handover certificates or contract completion dates, putting them in that file and re-rendering is the whole procedure.

Two different things would follow, and they are worth keeping apart. Inside Leura Falls, better dates will not change the conclusion: the detectable effect of 0.65 score points is set by how many treated sites exist and how variable they are, not by how well the intervention is dated. Better dates sharpen a null. Across the network, they do something else entirely — they turn 120 blocked assets in 59 catchments into a design that could find about 0.19 score points, which is Section 17.5.1’s whole argument and the reason dq:sqid-commissioning-dates sits where it does on the list.

One thing can be said now, whatever happens to the dates: the treatment catchments are where the works are warranted, on the imperviousness figure this book adopts. All four of the most impervious catchments monitored are in them — 86BKT (Kedumba Creek, 44%), 59BLA (Leura Falls Creek, 39%), 59.2BLA (Leura Falls Trib u/s Chelmsford Dr, 36%), 16GKT (Katoomba Creek, 34%). Those percentages are one method’s answer rather than a measurement, and the qualification matters more here than the ranking does. Imperviousness is estimated three ways in the data layer — from cadastral road corridors and address points, from land-use classes, and from census mesh blocks — and on these four catchments the three disagree by 16.7 to 24.2 percentage points. Chapter 11 Section 11.7 sets out the same disagreement for Glenbrook Lagoon and publishes the three-method span rather than the point estimate alone. The ordering is not method-free either: on land-use estimates 4 of the four most impervious catchments are treatment catchments, and on mesh-block estimates 1, with a different four sites. The street-and-address figure is adopted throughout the book because it is the only one of the three built from quantities measured inside each individual catchment rather than from a fraction assumed for a class — so what follows rests on that method being the right choice, not on the percentages being known to a point. Two of those four are the same creek, 59.2BLA being the relocated replacement for 59BLA. There is no exception inside the top four, and the most impervious catchment with no works recorded against it sits just outside: Woodford Creek Tributary (M6), 33% impervious, ranked 5 and 0.3 of a percentage point below the fourth — worth a look on the same argument. Targeting is otherwise well justified even though the effect is not yet measurable.