| Site | Waterway | Note |
|---|---|---|
| 09.2BBH | Centennial Glen Creek | New site for Centennial Glen - shifted 2024 due to safety issues & poor habitat at 09BBH |
| 59.2BLA | Leura Falls Trib u/s Chelmsford Dr | replacing 59BLA (Cumberland), due to private land access issues |
| 84NBX | Cripple Creek | Added 2021 |
| 85BKT | Kedumba Creek | Added 2022 for Kedumba Project |
| 86BKT | Kedumba Creek | Added 2022 for Kedumba Project |
| 87GHZ | Hazelbrook Creek tributary | New site added as control for Hazelbrook bifenthrin contamination incident 2023 |
| 88GHZ | Hazelbrook Creek tributary | New site added 2023 - impact site for bifenthrin contamination |
| P2-M5 | Bedford Tributary | Old Site not sampled after 2005 |
17 Evaluating the works
Blue Mountains City Council Healthy Waterways — statistical analysis
17.1 What this chapter is for
The question you actually asked is whether the stormwater treatment works have improved the creeks. This chapter answers it as far as the archive allows, which is not far, and says precisely what is missing.
Three things come out of it. Dates. No treatment asset in either database carries a construction date anyone can rely on, and one project report is the only source of real ones. Design. Even with perfect dates, one catchment cannot answer this question at any level of effort — a point worth settling before the next evaluation is commissioned rather than after. A null that carries no information, which is a different thing from a null that says the works did nothing.
Every intervention date below is read at render time from a single file. Those dates are provisional, they are the weakest part of the analysis, and replacing them is a one-file job — Section 17.7 sets out how.
17.2 The data behind this chapter
Each question and request below is set out again in What we need from you, with what it blocks, what an answer is worth and what it would cost you to find, ranked against every other ask in the report.
17.2.1 Questions only you can answer
What do these fields actually record — build, handover, capitalisation, load date or revaluation batch? Construction Date in Civica, AssetPit.Year, tblAIMOpenChannel.Datecreate and Effectdate, and Asset_Audit.OC_CommDate.
OC_CommDate is the largest populated date field we have — 204 of 254 treatment assets — and whether it is usable or junk turns entirely on this. Our reading is that it is a year-end capitalisation stamp: it has only eleven distinct values, every one falls on 29 or 30 June, 169 of them are 2016-06-29, and it disagrees with all five independently documented construction dates. But that is inference, and somebody in Assets can just tell us.
Refer to it as
dq:civica-date-semantics.
Given that a four-site comparison cannot detect anything, which do you want — evaluate the works on a response that varies less between sites, pool several works programs to reach the site count, or accept in advance that the evidence will be descriptive and say so?
Detecting half a score point needs roughly twelve treated and twelve control sites, sampled twice a year, for six years either side. At four per arm it is unreachable at any amount of sampling — the floor is 0.73 score points. The Leura Falls difference-in-differences returned +0.11 with a confidence interval from -0.36 to +0.58: a null that carries no information at all. Deciding this in advance is what stops a four-site study being run and then reported as “no effect”.
Refer to it as
dq:stormwater-eval-design-decision.
Would it be possible to flag a sample in the database when something has happened to the creek — a spill, a contamination, a fire, works in the channel — rather than leaving it to a free-text comment or to nothing at all?
The July 2012 bifenthrin contamination of Jamison Creek destroyed the macroinvertebrate community for three months and is completely invisible in the database. Twenty-four samples across the impact and recovery period carry no marker of any kind, and anyone analysing the Jamison series without knowing the history would read a catastrophic contamination as a poor run of monitoring results. The 2023 Hazelbrook incident is recorded, but only as two sentences in a site note. This one recovers nothing already lost - it is about the next incident, not this one.
Refer to it as
dq:flag-incidents-against-samples.
17.2.2 What would answer them
Where did the Leura Falls monitoring data end up? Sites E151 and E195-E199, six sites monthly from June 2014 to July 2017, plus rain-event inlet and outlet sampling of five treatment systems with measured pollutant reductions per system.
Not one E-code site exists in either database. This is the only dataset anywhere that could say whether an individual treatment device works — which is a different question from the network-level one the commissioning dates would unlock, and arguably the more useful of the two. It is your own data, recently collected, and somebody knows where it is.
Refer to it as
dq:leura-falls-monitoring.
When was each treatment asset built? The specific ask: a OneCouncil / CiAnywhere extract filtered to SWDR-GPTS and pit types 07, 08, 11 and 12, returning the FA450 financial-asset commissioning date — plus the DA and construction-certificate records, which have never been searched, and whoever populated “EConst. Date” on 28 SQIDs.
Your central question is whether the stormwater works have improved waterway health, and it cannot be answered because nothing reliably records when the works happened. This would take the evaluation from three usable comparison units to something near sixty, and the smallest change it could detect from more than half a score point to about a fifth of one. Be clear about what it does not do, though: dates make the question askable, not answerable. The monitoring design decides the rest, which is the next item but one.
Refer to it as
dq:sqid-commissioning-dates.
Which catchment — or at minimum which upstream drainage node — does each treatment asset serve? It exists nowhere except for the seven Leura Falls systems.
Without it, “downstream of a treatment device” cannot be defined at all, so no monitoring site can be paired to any asset and no asset can be evaluated regardless of how good its date is. This is the quieter half of the commissioning-date problem and it is just as binding.
Refer to it as
dq:sqid-served-catchment.
Do you still hold the two datasets that sit beside the 2012 Jamison Creek macroinvertebrate sampling — the OEH bifenthrin laboratory results (8 sites, 19 samples, July 2012) and Robert McCormack’s freshwater crayfish surveys (3 sites, 30 samples, July 2012 to April 2014)?
Your conference paper on the incident reports all three datasets together, but only the macroinvertebrate sampling reached the database we have. The bifenthrin results would give us an actual dose against an actual response, which nothing else in twenty-six years of macroinvertebrate monitoring can offer, and the crayfish surveys track the species the kill was noticed by. Both would turn a well-documented incident into a quantitative exposure-response series on one of your own creeks.
Refer to it as
dq:jamison-2012-companion-data.
Is there a list of the treatment assets that are not pits — rock-lined biofilters, raingardens, bioretention basins, sediment basins, trash racks — and GIS layers for them? They are largely invisible to the asset system.
We have a partial list already: sqid_dates.csv carries 315 treatment assets, including biofilters and rock-lined channels picked up from your 2021 Stormwater System Maintenance Program spreadsheet that are not in the GIS pit table. What is missing is a register that says the list is complete, so we cannot tell whether the gap is five assets or a hundred.
Refer to it as
dq:non-pit-assets.
Where would a decommissioning or major-maintenance date be recorded — is there an asset-disposal workflow in Civica/OneCouncil at all? The Disposed extract returns zero SQIDs, which we read as “none are captured” rather than “none have ever been removed”.
Same mechanism as the commissioning dates and smaller: without them we cannot tell whether an asset was actually operating during the monitoring period it is being credited with.
Refer to it as
dq:sqid-decommissioning.
Does a fine-grained stormwater sub-catchment layer exist, or an imperviousness layer to go with it? Data\Environment on the corporate GIS share is ACL-denied to us and its orphaned label layers point straight at sub-catchment polygons — could we be given read access to it?
It is what would let the drainage area served by an asset be defined at anything finer than the delineated site catchment, which is currently the only unit available and is far too coarse for a device serving a few streets.
Refer to it as
dq:stormwater-subcatchment-layer.
Would you ask Sydney Water for records dating the original Vale St wetland cells? They were built by Sydney Water in the 1990s rather than by Council, which is why they are not in your register.
One asset, but a load-bearing one: it is the intervention date for the one catchment with a defensible before-and-after case.
Refer to it as
dq:sydneywater-valest.
17.3 What the record holds
The methods document lists stormwater treatment projects in the Leura Falls, Jamison, Kedumba and Knapsack Creek catchments as one of the reasons sites were selected. The databases record almost nothing about them.
tblProjects in the water quality database lists Macro, RecWQ, Sediment, Jamison and Leura as projects, but only Macro and RecWQ are ever used: the Jamison and Leura catchment restoration projects have zero samples tagged. Project attribution therefore has to be inferred from site codes and dates, and no commissioning date for any treatment device exists in either database, in the methods document, or in Council’s public material.
There is one designed impact study in this archive, and nothing in the databases says so. In July 2012 more than a thousand dead Giant Spiny Crayfish turned up along two kilometres of Jamison Creek at Wentworth Falls. Bifenthrin had been over-applied at a residential development on 5 July, about 3 mm of rain fell on the 6th, and the kill was found the next day; the pesticide reached the creek 300 m away through a stormwater pit and pipe run, and two pest control operators were later convicted. All of this is in your own conference paper (St Lawrence et al. 2014) — none of it is in the monitoring database.
The sampling is there, though, and it is the closest thing in the archive to a designed impact study. Monthly sampling ran through the impact period at three sites: 61BWF 290 m upstream of where the pesticide entered, which is the control; 62BWF 120 m downstream and 23.2BWF 2.2 km downstream, which are the impact sites. The dates match the paper’s figure exactly. From 2013 the three drop back to the network’s routine cadence of twice a year, and the record does not stop there: the last visit is 3 April 2024. The signal is not subtle — across July, August and September 2012 the control holds 7, 13 and 16 families while the near site holds 1, 3 and 3.
It is not, however, a before–after–control–impact design. 61BWF and 62BWF were both first visited on 12 July 2012, five days after the kill was found: neither the control nor the near impact site has a single macroinvertebrate sample from before the over-application, and the water quality database — a separate record that could in principle have supplied one — holds none at either site either. The only pre-incident sampling anywhere in the set is at 23.2BWF, 2.2 km downstream: five samples between 3 April 2008 and 14 March 2012, the last of them four months before the spill. The Before × Control cell is therefore empty, and the interaction that defines a BACI — how much more the impact sites changed than the control did — cannot be estimated at all. What the archive actually holds is an after-only control–impact design with nineteen rounds of paired temporal replication, plus a before-and-after series at one of the two impact sites.
That is a limitation of the event, not of Council’s response. A pesticide spill is not scheduled, and both monitored sites exist because of it; no amount of diligence produces a pre-incident sample at a site established afterwards. The upstream reach is a spatial control — the standard and correct design for an impacts-and-recovery study of an unannounced contamination, which is exactly what St Lawrence et al. (2014) set out to do. The design is right for its purpose. It is simply not the design that would let this report difference out a catchment trend.
The control also drops out before the impact sites do, and anyone treating this as a control–impact comparison has to stop where it does. 61BWF was last sampled on 19 April 2021; 62BWF and 23.2BWF were sampled six more times after that, out to 3 April 2024. Those last six rounds have no contemporaneous upstream comparison, so they can only be read against the impact sites’ own history. The paired part of the design ends in 2021, and a recovery trajectory fitted past that year is an impact-site time series with the control quietly dropped, which is the inference the next section explains a control exists to prevent.
That matters here for two reasons beyond the incident itself. It is the only place in the 26 years of macroinvertebrate monitoring — 1998 to 2024, which is not the same span as the 27-year water quality record — where we can see what a known dose of a known contaminant does to these creeks and how long recovery takes. And the mechanism is the one this whole report keeps circling: the pesticide travelled from a property to a creek 300 m away through the stormwater network, which is what a connected, impervious catchment is. The paper’s own conclusion is about effective imperviousness, and it is the only published test of that mechanism on a Blue Mountains creek.
17.3.1 Could a before–after–control–impact design work at all?
A BACI analysis needs four things: samples at the treated site before the works, samples after, matched control sites monitored over the same period, and a known intervention date. The requirement is not a formality invented here — without a control, an impact cannot be separated from whatever else happened between the two periods, and without repeated sampling in each period the comparison has no error term to test against (Stewart-Oaten et al. 1986). Underwood (1994) goes further: a single control is not enough either, because a treated and a control site can drift apart for reasons unrelated to the works. That drift is exactly the variance component that dominates the power calculation in Section 17.5, so it is named here rather than there.
| Catchment | Site | Waterway | First year | Last year | Years sampled | Samples |
|---|---|---|---|---|---|---|
| Jamison | 23BWF | Jamison Creek | 1998 | 2007 | 7 | 15 |
| Jamison | 22BWF | Valley of the Waters | 2000 | 2013 | 10 | 19 |
| Jamison | 24BWF | Wentworth Falls Lake | 2006 | 2024 | 18 | 33 |
| Jamison | 23.2BWF | Jamison Creek | 2008 | 2024 | 16 | 31 |
| Jamison | 61BWF | Jamison Creek | 2012 | 2021 | 10 | 20 |
| Jamison | 62BWF | Jamison Creek | 2012 | 2024 | 13 | 25 |
| Jamison | 22.2BWF | Lillians Glen | 2015 | 2024 | 10 | 10 |
| Kedumba | 19BKT | Kedumba Creek | 1998 | 2024 | 23 | 38 |
| Kedumba | 16GKT | Katoomba Creek | 2000 | 2024 | 21 | 28 |
| Kedumba | P2-U8 | Katoomba Creek | 2000 | 2004 | 2 | 4 |
| Kedumba | 73BKT | Kedumba River | 2016 | 2023 | 5 | 6 |
| Kedumba | 85BKT | Kedumba Creek | 2022 | 2024 | 3 | 5 |
| Kedumba | 86BKT | Kedumba Creek | 2022 | 2024 | 3 | 5 |
| Leura Falls | 21BLA | Gordon Creek | 1998 | 2024 | 20 | 29 |
| Leura Falls | 20BLA | Leura Falls Creek | 1999 | 2024 | 22 | 46 |
| Leura Falls | 58BLA | Leura Falls Creek | 2011 | 2024 | 13 | 22 |
| Leura Falls | 59BLA | Leura Falls Creek | 2014 | 2021 | 8 | 16 |
| Leura Falls | 59.2BLA | Leura Falls Trib u/s Chelmsford Dr | 2023 | 2024 | 2 | 3 |
Three of the four named catchments, and the fourth is accounted for rather than dropped. Section 17.3 names Leura Falls, Jamison, Kedumba and Knapsack Creek; Table 17.2 carries the first three. Knapsack Creek is monitored, and monitored well — 27 samples at 1 site (50NEP) across 19 years from 2002 to 2024, a longer record than 12 of the 18 series in Table 17.2. It is left out because there is nothing in it to evaluate: the stormwater extract places no treatment asset at all in that site’s catchment group, while the only asset in the whole 315-asset extract whose location names Knapsack is a gross pollutant trap carrying nothing but a 1997 drainage-inventory date, which pre-dates the monitoring record and is excluded on the same rule as the rest of that block (Section 17.4). Read that as a hole in the asset record rather than as evidence that nothing was built there — the methods document says projects exist in the catchment, and the asset extract does not show them, which is one more thing dq:sqid-commissioning-dates would settle.
Kedumba: no. Sites 85BKT and 86BKT were, on your own note, “Added 2022 for Kedumba Project” (Table 17.1). They have 10 samples between them, all from 2022 onward. There is no before period, and no statistical method creates one. The only Kedumba sites with a long record (19BKT at Katoomba Cascades, 16GKT on Katoomba Creek) sit downstream of a catchment that has had continuous incremental works, so they integrate everything and isolate nothing.
Leura Falls: not as designed, but closest to feasible. 20BLA at Leura Cascades has 46 samples across 22 years from 1999 to 2024 — a genuine long series spanning any plausible works period. What is missing is a date and a control.
Jamison: confounded by a relocation and by the 2012 contamination. The long Jamison Creek series runs 23BWF (1998–2007) then 23.2BWF (2008–2024), a relocated pair that must not be spliced without comment — these are the paper’s ‘Armstrong St’ and ‘Weeping Rock’, 560 m and 2.2 km below the point where the bifenthrin entered. The join sits inside any plausible works window, so a step at the splice cannot be told apart from a step from the works, and the July 2012 contamination sits inside it too and would swamp either.
The design as recorded therefore cannot demonstrate whether the works changed waterway health — not because the sites are poor, several are excellent long series, but because the two things that turn a monitoring series into an experiment are absent: when each device was commissioned, and which sites are untreated controls. Two partial fixes have since arrived: the stormwater asset extract supplies dates of a sort (Section 17.4), and your own 2018 project report supplies real construction windows for seven systems in one catchment (Section 17.6). A narrow evaluation is now possible. It returns nothing, and the reason it returns nothing is that it was never capable of returning anything smaller than an improvement larger than any plausible treatment effect.
17.4 What the asset register can and cannot date
Your stormwater extract carries 315 treatment assets with a date of some kind against them. Almost none of those dates means what an evaluation needs it to mean.
| Stage | Assets |
|---|---|
| Treatment assets with any date at all | 315 |
| … whose date plausibly means “built” | 77 |
| … of those, high confidence | 24 |
| … whose date is an inspection (upper bound only) | 227 |
| … with no usable date | 11 |
| In the 1997 drainage inventory (pre-date the record entirely) | 74 |
| With a coordinate, so placeable in a catchment | 265 |
Three features of Table 17.3 decide what can be attempted.
Most dates are inspections. 227 of the 315 assets are dated by the first or only time somebody inspected them, which bounds construction from above and nothing more. Treating one as an intervention date would place the works at whatever moment record-keeping happened to start — for many assets 2021 or 2022, years after they were built. Three of the seven systems below carry 2021 inspection dates against construction your own project report puts in 2016.
A large block pre-dates the record entirely. 74 pit-based assets carry a 1997 drainage-inventory date. The macroinvertebrate record begins after that, so they have no before period and are excluded throughout. This is also what removes Knapsack Creek from the BACI design. The extract places 0 treatment assets in the catchment group of 50NEP, and the only asset in the whole extract whose location string names Knapsack — a gross pollutant trap by Knapsack Park at Glenbrook — is in this 1997 block, dated by an inspection and flagged low confidence. So the fourth catchment the methods document names has nothing datable in it to evaluate.
The accounting dates are not build dates. The Civica extract offers a commissioning date for all 204 treatment assets in the asset audit, and every one of the 204 falls on 29 or 30 June — financial year end — with 169 in 2016, the year of an asset revaluation. They are capitalisation dates: the Glenbrook Lagoon gross pollutant traps are stamped 1 January 2015 against notification forms showing construction in April and May 2014. Nothing here uses them.
17.4.1 The size of the prize
| Stage | Assets | Site pairings | Catchment groups |
|---|---|---|---|
| Treatment assets placed in a monitored catchment | 214 | 509 | 62 |
| … that do not pre-date the monitoring record | 169 | 413 | 59 |
| … of those, dated to a construction event | 49 | 118 | 27 |
| … dated only by an inspection (upper bound) | 109 | 268 | 49 |
| … carrying no date at all | 11 | 27 | 15 |
Table 17.4 is the most important thing in this chapter. 120 of the 169 candidate treatment assets — 71% — are blocked purely by date quality. Not by geography, not by monitoring coverage, not by the works being too small: by nobody having written down when they were built. Counted as site pairings instead — the unit an evaluation actually works in, since one asset upstream of several monitoring sites offers several comparisons — it is 295 of 413, or 71%: the same share either way, so nothing here turns on which unit you count in. Those assets sit in 59 catchment groups, against the 3 treated sites the Leura Falls evaluation actually has, and there are a further 58 catchment groups in the network with no treatment asset at all to serve as controls.
Installation dates are what the evaluation question Council asked for is waiting on (dq:sqid-commissioning-dates), and Section 17.5 says what they would buy, in the units the design is measured in. They are not the report’s single highest-value request: chapter 2’s master list rates 14 asks transformative and ranks this one 12 of 107 once effort is set against value. This chapter’s own dq:leura-falls-monitoring ranks above it, and it should — dates make the network question askable, while the Leura Falls data is the only thing in existence that could evaluate a single device.
17.5 How much monitoring an evaluation needs
| Component | Treatment catchments | Whole network |
|---|---|---|
| Between sites | 0.57 | 0.68 |
| Shared between periods (differenced out by a control arm) | 0.22 | 0.38 |
| Site x period — how differently sites move (tau) | 0.18 | 0.10 |
| Site x year — how much one site moves year to year (sigma_year) | 0.29 | 0.33 |
| Within site and year — sampling noise (sigma) | 0.47 | 0.45 |
| Sites per arm | Power to detect 0.50 | 95% CI | Smallest detectable change | Floor (unlimited sampling) |
|---|---|---|---|---|
| 4 | 40% | 35-44% | 0.86 | 0.73 |
| 8 | 75% | 71-79% | 0.54 | 0.46 |
| 10 | 79% | 76-83% | 0.48 | 0.41 |
| 12 | 91% | 89-94% | 0.43 | 0.37 |
| 14 | 96% | 93-97% | 0.40 | 0.34 |
| 16 | 98% | 97-99% | 0.37 | 0.31 |
Two sources of variation matter here, and counting only one of them is the commonest way to under-resource a monitoring program.
The first is sampling noise: a single sample at a single site is a noisy measure of that site’s condition, standard deviation 0.47 units (Table 17.5). More sampling fixes that. The second is that sites do not all change by the same amount between two periods, even without any treatment — a creek recovers from a fire, another silts up, a third changes for no reason anyone recorded. Measured on these very sites the standard deviation of that movement is 0.18 units, and the whole-network estimate of 0.10 says it is not a quirk of the three treatment catchments. Because a difference-in-differences contrast is a difference of site-level changes, this second component is divided by the number of sites, not the number of samples. No amount of extra sampling reduces it.
There is a third, and leaving it out is what makes a power calculation flatter itself. A single site also moves from one year to the next — standard deviation 0.29 units here, 0.33 network-wide — and two samples taken in the same year share that movement. Counting them as two independent draws is the arithmetic that turns twelve samples over six years into twelve independent measurements when they are nearer to six. That component is divided by the number of years, so it is reduced by a longer record and not by a busier one.
What that costs:
- Four treated and four control sites, sampled twice a year for six years either side, has only a 40% chance of detecting an improvement of half a score point that genuinely occurred. The change it could reliably detect is about 0.86 score points.
- At four sites per arm, half a score point is out of reach at any amount of sampling: the floor is 0.73 score points.
- Detecting half a score point needs about 12 treated and 12 control sites, each sampled twice a year for about six years either side. That reaches 91% power, and the closed form agrees at 0.43 score points.
- A change of a quarter of a score point is not worth designing for: it needs at least 25 sites per arm on the floor alone.
Judgement call 7 is answered by Table 17.6, and the answer is unwelcome: one catchment cannot answer this question at any level of effort. If 12 sites per arm is beyond reach — and for most councils it will be — the conclusion is not to run a smaller version of the same design. It is that a difference-in-differences on the health score is the wrong instrument, and you would do better to evaluate the works on a response that varies less between sites (a targeted water quality or physical habitat measure at the outfall), or to accept in advance that the evidence will be descriptive rather than statistical. Designing a four-site study and reporting its null as “no effect” is the worst of the three.
17.5.1 What fixing the dates would buy
| Scenario | Treated units | Control units | Smallest detectable, 12 samples each side | Smallest detectable, 6 samples each side | Floor (unlimited sampling) |
|---|---|---|---|---|---|
| Now — the seven Leura Falls systems | 3 | 9 | 0.75 | 0.85 | 0.64 |
| Catchments with an asset already dated to a construction event | 27 | 58 | 0.24 | 0.27 | 0.20 |
| Every candidate asset dated | 59 | 58 | 0.19 | 0.21 | 0.16 |
So dating the assets moves the evaluation from a design that cannot see 0.65 score points to one that could find about 0.19 score points, with a floor of 0.16 that no extra sampling removes (Table 17.7). That is about a fifth of a score point, and it is the region where a real catchment-scale treatment effect might plausibly live.
Two honest riders, because this is the number the whole data request rests on.
It is a ceiling, not a forecast. It assumes every one of the 59 catchments is datable, that each has twelve edge samples either side of its own intervention, and that the intervention dates are far enough apart not to collide with the before periods of the others. None of those is guaranteed; the real figure will be worse.
The brief’s range describes the floor, not the evaluation. The floor of 0.16 does sit inside the “0.1 to 0.2” the brief claimed, so on magnitude the brief is close. What it gets wrong is what the number is. 0.16 is an asymptote: it is set by tau alone, it is what this design converges on as sampling goes to infinity, and there is no budget, on this response and with this many catchments, that betters it. The figure a real program would achieve is the one above it — 0.19 at twelve samples a side — and the rider above makes even that a ceiling rather than a forecast. So quote 0.19, name 0.16 as the floor beneath it, and do not quote a range beginning with 0.1: nothing in this design reaches 0.1.
17.6 The seven Leura Falls systems
Between September 2015 and June 2017 seven stormwater treatment systems were built in the Leura Falls Creek catchment under the Leura Falls Catchment Improvement Project (Table 17.8). They are the only treatment assets anywhere in the network with a recorded served catchment, and the project’s own final report (Blue Mountains City Council 2018) states a construction window for each one. That report is the source of every date used here.
Constructed is the end of the construction window the report gives, because a system treats nothing until it is finished. The report writes “constructed around …” for every one of the seven, so none of these dates is better than month precision. The Vale St catchment figure covers both flowlines. These dates, and only these, are what the evaluation below runs on; they live in R/data-layer/stormwater_intervention_dates.csv and can be replaced without touching any code.
| System | Town | Catchment (ha) | Impervious | Constructed | Confidence |
|---|---|---|---|---|---|
| Vale St southern flowline | Katoomba | 24 | 55% | Apr 2016 | medium |
| Vale St northern flowline | Katoomba | — | — | Jun 2017 | medium |
| Jersey Ave | Leura | 4 | 60% | Mar 2016 | medium |
| Craigend St | Leura | 1 | 60% | Mar 2016 | medium |
| Murray St | Leura | 5 | 70% | Mar 2016 | medium |
| Kanimbla St | Katoomba | 11 | 30% | Sep 2015 | medium |
| Commonwealth St | Leura | 6 | 50% | Oct 2016 | medium |
17.6.1 Which monitoring site is downstream of which system
A treatment system only counts as an intervention at a monitoring site if the ground it treats actually drains to that site. That is a drainage question, not a distance question, and it is settled here against the DEM-delineated catchments rather than by proximity. Six of the seven systems have no coordinate in any asset register — the rock-lined biofilters and raingardens are largely invisible to the asset system — so each system is represented by its road corridor, taken from the AssetRoadSurface layer, and the evidence published for each pair is how much of that corridor falls inside the site’s catchment.
| System | Monitoring site | Site catchment (ha) | Corridor in catchment |
|---|---|---|---|
| Commonwealth St | 20BLA | 215 | 137 m of 137 m (100%) |
| Craigend St | 20BLA | 215 | 239 m of 2058 m (12%) |
| Craigend St | 21BLA | 175 | 1190 m of 2058 m (58%) |
| Jersey Ave | 20BLA | 215 | 586 m of 790 m (74%) |
| Kanimbla St | 20BLA | 215 | 706 m of 706 m (100%) |
| Kanimbla St | 58BLA | 73 | 577 m of 706 m (82%) |
| Murray St | 20BLA | 215 | 344 m of 344 m (100%) |
| Vale St northern flowline | 20BLA | 215 | 395 m of 556 m (71%) |
| Vale St northern flowline | 59.2BLA | 66 | 395 m of 556 m (71%) |
| Vale St northern flowline | 59BLA | 39 | 395 m of 556 m (71%) |
| Vale St southern flowline | 20BLA | 215 | 395 m of 556 m (71%) |
| Vale St southern flowline | 59.2BLA | 66 | 395 m of 556 m (71%) |
| Vale St southern flowline | 59BLA | 39 | 395 m of 556 m (71%) |
Two consequences fall straight out of Table 17.9, and both cost the evaluation something.
The obvious control is not a control. Gordon Creek (21BLA) is the site you would otherwise reach for: a long series in the same town, outside the Leura Falls Creek catchment. But 58% of the Craigend St road corridor drains to Gordon Creek, so Gordon Creek is downstream of a treatment system and belongs in the impact arm, not the control arm. It is a poor member of either: Craigend St treats 1.0 ha of a 175 ha catchment, or 0.6% of it (Table 17.10). It is dropped from both arms.
The Commonwealth St system is downstream of the Commonwealth St site. Site 58BLA is named for Commonwealth St, but the Commonwealth St road corridor lies outside 58BLA’s catchment. 58BLA is upstream of that system and cannot be used to evaluate it. It is downstream of Kanimbla St, and is used for that.
17.6.2 The result
| Site | Before | During | After | Verdict |
|---|---|---|---|---|
| 59BLA | 4 | 3 | 9 | usable |
| 59.2BLA | 0 | 0 | 3 | excluded — too few samples before |
| 20BLA | 20 | 3 | 14 | usable |
| 58BLA | 6 | 4 | 12 | usable |
| 21BLA | 13 | 2 | 7 | excluded — 0.6% of catchment treated |
Of the 5 sites with a treatment system upstream, 3 can carry a before-and-after comparison at all: 59BLA, 20BLA, 58BLA (Table 17.11). 59.2BLA, the relocated replacement for 59BLA, has no samples before the works and can never contribute — the same defect as the Kedumba sites, for the same reason. Three impact sites, against the 12 per arm Section 17.5 says an evaluation that could detect half a score point needs.
The difference-in-differences estimate of the effect of the Leura Falls systems on the health score is +0.11 score points (95% CI -0.36 to +0.58), from 213 edge samples at 3 treated and 9 control sites, fitted with site, year and site-by-period random effects. One rating class is 1.00 score points wide. That is a null result, and the interval is wide enough to contain both a deterioration of 0.36 score points and an improvement of 0.58.
The detectable effect size is the finding. At 80% power and 5% significance, this design could not have detected an improvement smaller than 0.65 score points — 65% of a score point, a jump most of the way from one rating class to the next. No stormwater treatment program anywhere would be expected to deliver that at the catchment scale. The observed estimate is 17% of it. The same closed form as Table 17.7, evaluated at this design’s own record rather than at the round six and twelve tabulated there — 8.9 samples per site and period, over 6.6 years — is more pessimistic again — 0.78 score points, of which 0.63 is floor — and both are more than half a score point, which is the only comparison that matters.
So the honest reading of the null is not “the works did not improve the creek” but “this design could not have told you either way” — and Section 17.5 said so in advance, before any of these data were fitted.
17.6.3 What the null is and is not confounded with
| Period | hydrolab | aquaread | aquatroll |
|---|---|---|---|
| before | 71 | 0 | 0 |
| after | 0 | 24 | 75 |
The intervention window sits directly on the first of the two dissolved-oxygen instrument steps Section 3.7.2 identifies. Period and instrument era are the same variable here (Table 17.12), and no model can separate them from a single arm. This is exactly why the design needs controls rather than a before-and-after: the control sites went through the same probe replacements on the same dates, so the period-by-role interaction differences the instrument out even though the period main effect is uninterpretable. It is also why the era term cannot be carried as a covariate here the way Section 6.4 carries it — there is nothing left for it to estimate.
The health score, which carries the primary result, is not exposed at all: it is built from families identified under a microscope, not from a probe. Dissolved oxygen, pH, electrical conductivity, temperature and turbidity are fully exposed and are only interpretable through the control arm. Alkalinity, phosphate, nitrate and faecal coliforms come from test kits and are untouched by the probe changes. One of those four carries a different unknown, and it is about the analyte rather than the instrument: nothing in the record says whether phosphate is reported as PO4 or as P, and the two differ by a factor of three (Section 7.5). That is not a further confound here — it is one constant factor on one column, so it cancels out of a before-and-after contrast taken within it — but it is why nitrate-N declares its species in its name and phosphate does not.
Two further confounds land inside the after period, and neither can be differenced away by a control site outside the catchment, because both are specific to Leura Falls Creek:
- A raw sewage leak ran into Leura Falls Creek for at least 22 months, from April 2016 to February 2018 (Blue Mountains City Council 2018). It was reported to Sydney Water and the EPA in April 2016; the source was not located until your own officers found it in February 2018. It entered the creek downstream of the upper monitoring sites but upstream of Leura Cascades (20BLA) and the Commonwealth St site, two of the three impact sites. The leak begins one month after four of the seven systems were finished and runs through the whole early after period.
- A landscaping supplies business was discharging enough phosphorus and suspended solids into the upper catchment for a Prevention Notice to issue in September 2017 (Blue Mountains City Council 2018).
Both push the after period in the opposite direction to the treatment. A null in the presence of a 22-month sewage leak is not evidence that the works did nothing; it is evidence that the design cannot see past a much larger signal. The 2020-onward wet years (Section 4.4) overlap the after window too, though that one the control arm does difference out.
Your own 2018 report reached the same conclusion from the same catchment. It records “no clear improvement in sub-catchment-wide downstream water quality after construction of stormwater treatments”, and “no clear evidence of improvement in waterway health ratings” (Blue Mountains City Council 2018). What is added here is a control arm and a stated detectable effect size, which is the part that turns a disappointment into a design specification.
| Parameter | n | Treated sites | Effect | 95% CI | Smallest detectable |
|---|---|---|---|---|---|
| Dissolved oxygen (% saturation) | 152 | 2 | -11.48 | -40.37 to +17.42 | 39.90 |
| Dissolved oxygen (mg/L) | 159 | 2 | -2.00 | -5.20 to +1.20 | 4.42 |
| pH | 164 | 2 | +0.11 | -0.49 to +0.72 | 0.84 |
Only 2 treated sites survive the water quality coverage requirement (Table 17.13), against 3 for the health score, so these are weaker still.
17.7 If better dates arrive
Every intervention date in this chapter is read at render time from stormwater_intervention_dates.csv, in R/data-layer/. Nothing downstream of that file contains a date; the cached spatial matching holds none either, so it does not need rebuilding when one changes. Replacing a date and re-rendering moves every number, table and sentence above that depends on it.
That is deliberate, because the dates are provisional and should be read that way. The seven systems are dated from a report that writes “constructed around” in front of each one, and the rest of the register is inspections and capitalisation dates. If Healthy Waterways or Assets hold commissioning records, handover certificates or contract completion dates, putting them in that file and re-rendering is the whole procedure.
Two different things would follow, and they are worth keeping apart. Inside Leura Falls, better dates will not change the conclusion: the detectable effect of 0.65 score points is set by how many treated sites exist and how variable they are, not by how well the intervention is dated. Better dates sharpen a null. Across the network, they do something else entirely — they turn 120 blocked assets in 59 catchments into a design that could find about 0.19 score points, which is Section 17.5.1’s whole argument and the reason dq:sqid-commissioning-dates sits where it does on the list.
One thing can be said now, whatever happens to the dates: the treatment catchments are where the works are warranted, on the imperviousness figure this book adopts. All four of the most impervious catchments monitored are in them — 86BKT (Kedumba Creek, 44%), 59BLA (Leura Falls Creek, 39%), 59.2BLA (Leura Falls Trib u/s Chelmsford Dr, 36%), 16GKT (Katoomba Creek, 34%). Those percentages are one method’s answer rather than a measurement, and the qualification matters more here than the ranking does. Imperviousness is estimated three ways in the data layer — from cadastral road corridors and address points, from land-use classes, and from census mesh blocks — and on these four catchments the three disagree by 16.7 to 24.2 percentage points. Chapter 11 Section 11.7 sets out the same disagreement for Glenbrook Lagoon and publishes the three-method span rather than the point estimate alone. The ordering is not method-free either: on land-use estimates 4 of the four most impervious catchments are treatment catchments, and on mesh-block estimates 1, with a different four sites. The street-and-address figure is adopted throughout the book because it is the only one of the three built from quantities measured inside each individual catchment rather than from a fraction assumed for a class — so what follows rests on that method being the right choice, not on the percentages being known to a point. Two of those four are the same creek, 59.2BLA being the relocated replacement for 59BLA. There is no exception inside the top four, and the most impervious catchment with no works recorded against it sits just outside: Woodford Creek Tributary (M6), 33% impervious, ranked 5 and 0.3 of a percentage point below the fourth — worth a look on the same argument. Targeting is otherwise well justified even though the effect is not yet measurable.