| # | Suggestion | Strength | Evidence in | Effort |
|---|---|---|---|---|
| R1 | Rate from three years, not one sample | Compels | ch 13 | None |
| R1b | Three rating classes, or score + interval | Supports | ch 13 | None |
| R2 | Rarefy the count factors; drop SIGNAL-SF | Supports | ch 15 | An afternoon |
| R3 | Score once, between two anchors | Supports | ch 15 | An afternoon |
| R4 | Reset the percentile anchors | Client ask | ch 14 | Low |
| R4a | Do not do R4 without R3 | Compels | ch 14, 15 | None |
| R4b | Record window, burn extent and severity | Judgement | ch 14 | Low |
| R5 | Water quality as its own report card | Supports | ch 6, 7, 9 | Moderate |
| R6 | Always publish the four factors | Supports | ch 5, 13 | None |
| R7 | Flag the sensitive rare families | Supports | ch 10, 15, 18 | Low |
| R8 | Split the anchors by distance from source | Judgement | ch 15 | Low |
| R9 | Generate the band table from a script | Compels | ch 14 | Low |
| R10 | Riparian and geomorphic third panel | Client ask | ch 9 | Moderate |
| R11 | A before-and-after study of your own works | Judgement | ch 17 | High |
| R12 | Test any future rating on the six criteria | Compels | ch 16 | Low |
| R13 | Publish the count of withheld ratings | Compels | ch 15 | None |
| R14 | A wetland ranking, and a second reference wetland | Supports | ch 11 | Low |
19 Conclusions
Blue Mountains City Council Healthy Waterways — statistical analysis
19.1 What this chapter is for
The preface answers the three questions you asked. This chapter does the harder thing: it pulls the analysis chapters together into one account of what is happening in Blue Mountains creeks, and says what we think you might do about it, with the reasoning, the cost and the consequence attached to each suggestion.
Two of those things are here in full — what the record supports concluding about the creeks (Section 19.2), and what to do about the rating system (Section 19.3), which is the largest single body of work in the report. Three belong with the evidence for them and are only pointed at from here: what to change in the monitoring program is chapter 2, Section 2.9; whether to create a protection tier for the creeks holding the rare animals is chapter 10, Section 10.6; and what to analyse next is chapter 18, Section 18.4. What we could not close is Section 19.5, and the full version of that is chapter 2’s Section 2.3.
Nothing here is new evidence. Every figure quoted is developed in an earlier chapter and cross-referenced to it, and where it comes from a fitted model it is read from that model rather than retyped.
19.2 What we now know about Blue Mountains waterway health
19.2.1 The recovery is real, and it ended about 2014
Averaged across the whole record, the published waterway health score rises by 0.40 points per decade on the 0-to-5 scale (95% confidence interval 0.30 to 0.51) — about two fifths of a rating class each decade. Across the twenty-six years the fitted curve rises 0.88 score points, most of a rating class, from the lower part of Fair into Good (chapter 5).
The rise does not run to the end of the record. Chapter 5 estimates where it stops rather than assuming a round number, and puts the changepoint at 2014.4. That date is well pinned — drop any single year of the record and it moves by less than half a year — but the window around it is wide once a whole field round is allowed to share its weather and its crew: 95% CI 2010.5 to 2017.25. That interval is a likelihood profile, a separate fit from the one the estimate comes from, because the estimator itself cannot carry a year effect; the profile puts the change at 2014.5, 0.1 of a year away and in the same calendar year, and 2014 is what the record is split at. Before the changepoint the score rises 0.58 points per decade (0.40 to 0.76); after it the estimate is −0.12 (−0.46 to 0.22) — the first three rows of Table 5.3 — an interval still wide enough to hold a decline of 0.46 points per decade, or an improvement of 0.22, so it is an absence of detected change and not a demonstration of stability. The two sides are not the same size, and this window’s wide side is the decline: read as movement “in either direction” it would overstate the improvement the record could be hiding by more than double.
Three separate checks say the rise is not an accident of the record.
- It is not the drought breaking. The record begins inside the Millennium Drought and ends after the wettest years in the series. Adding the 12-month drought index and antecedent rainfall to the model moves the trend by an amount of no practical consequence (chapter 5).
- It is not an artefact of which creeks were visited. The monitoring network changed shape completely over the period — no site was sampled in every year from 2006 — but the trend survives a site random effect, a restriction to the fourteen creeks monitored every year since 2012, and a restriction to your own core reporting sites (chapter 5).
- It shows up in data independent of the animals. Water quality measurements fall inside your desirable ranges more often than they used to, 65% over 2022–2025 against 54% over 2002–2005 across 11,178 measurements, and the improvement is indistinguishable in urban catchments and reference ones (chapter 6). Two of the parameters carrying it are the two whose measurement changed most, so that caveat should travel with the figure — but the improvement survives removing them. The animals and the chemistry are different measurements taken by different means, and they agree on the direction.
The plateau is much the weaker of the two findings, and worth saying carefully: what the record establishes is that no movement has been detected in a decade, not that none has occurred. Restricting to 2010 onward narrows the interval to −0.13 to 0.25 score points per decade, which is better but still wide enough to hold an improvement of 0.25 points a decade, or a decline of 0.13 (chapter 5). On this window the wide side is the improvement, the opposite way round from the post-changepoint interval above; neither is symmetric and neither should be quoted as one number. The practical consequence is that the annual snapshot is not going to move, on the current index and at the current sampling intensity, whatever the creeks do — and the next section is about why that is partly a property of the index rather than of the creeks.
19.2.2 The recovery is a trade, not a gain — and since 2010, a loss
This is the finding that should replace the current public message.
Blue Mountains creeks are not carrying more kinds of animal than they did in 1998. Standardised to a fixed number of individuals counted, family richness shows no net change across twenty-six years (0.05 families per decade, -0.39 to 0.50; chapter 5) — and that null is the average of a rise and a fall, not the absence of both. Over 2010–2024 the same measure falls -0.85 (-1.65 to -0.05), on an interval that excludes zero, and the raw count falls with it, so the recent loss is in the creeks rather than in the standardising. What has changed alongside it is which animals are there. Over the record the average sample’s share of individuals belonging to pollution-tolerant families fell from 42% to 17%, and the share belonging to sensitive families rose from 46% to 64% (chapter 8). Pooling every animal counted rather than averaging over samples gives 40% to 15% — the same story either way. These are proportions, so they are unaffected by how many animals the laboratory picked from each sample — which makes them the most trustworthy quantities in the entire dataset.
The apparent decline in family counts since 2010 is part of the same story rather than a contradiction of it. Of the families lost per sample since 2010, 60% are graded tolerant, on a sampling interval of 40% to 106% (600 creek-cluster bootstrap resamples, chapter 8). Two further choices move it and neither is sampling error: moving the tolerance boundary by one SIGNAL 2 grade sweeps the share over 37% to 81%, and how large a loss it is a share of turns on starting the clock at 2010. Sensitive families have not declined at all under any of those choices (chapter 8). The creeks are shedding the animals that indicate poor water.
The upper limit above 100% is not a slip and should not be read as one. In 3.3% of the resamples the tolerant families fell by more than the whole loss, because the intermediate and sensitive families gained over the same window — a share above one, not an impossible one. The loss it is a share of was negative in all 600 resamples, so the decline itself is not in doubt at any resample; only how much of it to attribute to tolerant families is.
The individual families bear this out. Of the 75 families with enough records to test, 17 have a changed occupancy: 12 becoming more widespread, overwhelmingly mayflies, stoneflies and caddisflies (Table 8.8), and 5 becoming less widespread, of which the clearest are mosquitoes and water striders (chapter 8, Table 8.9).
And the creeks have kept their individual identities. Biotic homogenisation — every creek converging on the same short list of survivors, the usual signature of urban catchments worldwide (McKinney 2006) — is not detectable here, and that is a bound rather than an absence. Average between-creek dissimilarity is flat over the record, and the test can rule out a change larger than about 6.4% of the observed level of dissimilarity; a change smaller than that it cannot rule out, and nothing here shows there is none (chapter 8, Section 8.7). What can be said is that these creeks are still distinguishable from one another as far as these data can see — which could not have been assumed, because urbanisation is the most consistently reported driver of homogenisation there is and these are urbanising catchments — and that protecting them individually is therefore still worth doing.
The sentence to publish. Twenty-five years ago two in five animals in a Blue Mountains creek were pollution-tolerant species; today it is fewer than one in five. It is true, it is monotonic, no factor in the current rating reports it, and it is effort-robust: rarefying every sample to a constant 50 animals reproduces it (40% to 16%), and to 100 animals reproduces it again (40% to 14%) (chapter 8).
19.2.3 The recovery is measured badly by the current index
Three properties of the published rating system get in the way of seeing any of the above.
It averages away its own signal. The four factors do not move together. Over the whole record one of them is largely a measure of laboratory effort, one is a real compositional signal, and one — SIGNAL-SF — barely responds at all, rising a quarter as far as SIGNAL 2 does on the same samples. Since the changepoint the family count has clearly fallen while SIGNAL-SF has drifted up (chapter 5). Averaging four parts of that kind gives each equal weight regardless of what it is measuring, and the composite reads flat. A composite that hides its components cannot explain itself.
Two of its four factors partly measure laboratory effort. The number of individuals in a typical sample rose from 92 over 1998–2009 to 167 from 2014, and counts of families rise mechanically with the number of animals counted. About 31% of the apparent improvement in the health score disappears once sample abundance is controlled for, and the whole of the apparent gain in family richness does (chapter 5). Across the interquartile range of sample abundance the published score moves 0.41 points (0.36 to 0.46) — 41% of a full rating class — with no change whatever in the creek, and 0.60 points between one creek and another.
Its sensitivity factor sees about a quarter of what a full sensitivity index sees, and nothing that the full index does not. SIGNAL-SF has no grade for 91 of the 233 recorded taxon names, but the missing grades are the smaller problem. The grades themselves are compressed at the tolerant end. Across every family tested, SIGNAL 2 explains 22% of the variation in how fast a family’s occupancy is changing (0.130 in log-odds per decade per grade, p = 2.2e-05); SIGNAL-SF explains 5.7% (0.077 per grade, p = 0.039); and put both in one model, SIGNAL 2 keeps its coefficient while SIGNAL-SF’s falls to −0.05 and stops being distinguishable from zero (p = 0.28) (chapter 8).
SIGNAL-SF grades mosquito larvae 6 and diving beetles 7 — scores that describe clean-water animals — so the disappearance of exactly the families that are disappearing barely registers. SIGNAL-SF is a noisier version of a scale that works.
Put together: the index is built so that the largest, most reliable signal in the data is the one it reports least well.
19.2.4 Why the creeks differ from one another
Where a creek sits on the scale is a different question from whether it is changing, and the report answers it better than it answers the trend question.
Creek identity dominates. More than forty per cent of what distinguishes one sample from another is simply which creek it came from (chapter 8). Any analysis that does not model creek identity is explaining the wrong variance.
Of the measurable drivers, physical habitat ranks first — ahead of water quality, ahead of the passage of time, and far ahead of imperviousness (chapter 9). Read that as an ordering, because that is all it is. The ordering is now tested rather than assumed: chapter 9 rebuilds each block out of random columns of the same width and asks what such a block earns by itself, and the ordering survives that test. The sizes do not. More than half of what the physical habitat and water quality blocks each appear to explain uniquely is what same-width noise earns, and the noise floor rises with the number of variables in a block — so two blocks of different width cannot be compared on that number at all, and none of these shares should be quoted as an effect size. The data for it came from the shading, substrate and vegetation observations your field officers have written on the water quality field sheet since 2006, which no analysis had previously used. Adding physical habitat and water quality to chapter 8’s four blocks raises explained compositional variation from 19% to 31% on identical samples, and means measured site variables now account for about three-quarters of everything that separates one creek from another.
Those two percentages are an upper bound and chapter 9 says so plainly. The Bray–Curtis distance the ordination uses is not Euclidean, so a quarter of the total eigenvalue mass is negative and every percentage in that section is divided by a denominator smaller than the real variation. Refitted on a metric version of the same distance the rise is 10.6% to 18.3%, and the share of chapter 8’s unexplained gap that the two new blocks close is 8.6% rather than 15% (Table 9.7). The ordering of the blocks is unchanged on either scale, and the ordering is what the suggestions rest on.
Of the drivers you can act on, catchment imperviousness ranks first, and riparian shading is the strongest lever at the reach. Between creeks imperviousness explains 13% of compositional difference on its own — first of all 27 candidates in Table 9.9 and the first variable forward selection picks — and shading 5.0%, which is thirteenth on its own and survives selection. Once the whole model is assembled those margins compress and the order changes: what imperviousness contributes that no other variable does is 2.8%, fourth of the ten variables selection retains. Only distance from source is clearly ahead of it: the second-placed variable leads imperviousness by 0.26 of a percentage point, which is not a gap, and the order inside that group has already inverted twice between builds. It ranks first on what it explains outright, which is the claim to lean on; it does not carry the largest unique contribution and no sentence in this report should say that it does. Shading is the strongest lever at the reach for a different reason again: it is the one variable in the list a revegetation program changes inside a decade (chapter 9), not the one with the largest share of anything. Almost everything else in the ranking — position in the drainage network, catchment area, altitude, rainfall — is fixed and cannot be managed. Two warnings come with that list. Moss cover is worth recording and not worth acting on. Fitted on its own it explains 5.0%, twelfth of the 27, and about 46% of that is shared with riparian shading, the substrate and the water chemistry, which already knew it. What is left over is 2.7%, fifth of the ten retained — a real place in the model, and ahead of shading’s own 1.2%. Read that as what you would expect of something that grows where the conditions the sensitive animals need already exist. The size of that contribution is not the point and has moved in both directions between builds; the point is that moss follows the conditions rather than setting them. It is a cheap, field-visible summary of those conditions and quick to record, so keep recording it — but read it as an indicator, not a lever, and do not plant moss. And neither fine sediment nor recorded sedimentation ranks between creeks at all — that is a statement about what these field variables separate, not about whether sediment matters, and fine sediment is the strongest single predictor of sensitive rare-family richness (chapter 10).
Imperviousness works through habitat and chemistry rather than around them (chapter 9): in the block partition fitted over all 1,230 samples, its unique share of compositional variation falls from 3.6% to 0.5% once habitat and chemistry enter the model, against 6.1% when that block is fitted alone. That last share is a different quantity from the 13% above and does not contradict it: the ranking is fitted on the 74 between-creek site means and this partition on individual samples, so imperviousness accounts for more of what separates one creek from another than of the variation across the record as a whole. That is the mechanistic reason source control and stormwater retrofit are the right levers rather than creek works — the pathway runs through the reach, so the reach is where the effect shows up.
19.2.5 What has not changed, and what is being lost
The gap between urban and reference creeks has not closed. All three disturbance tiers improved over the record, and the urban trend is statistically indistinguishable from the reference trend; if anything the most impervious catchments improved slightly less (chapter 5). The same conclusion arrives independently from the water quality exceedance data (chapter 6). Both should be read as “no difference detected” rather than “the tiers improved at the same rate”: chapter 7’s tier comparison puts the urban-against-reference difference anywhere between 0.82 and 1.32 times the reference odds of an exceedance per decade, so a gap that was closing, or widening, by a good deal would still have come back as “no difference”. If the expectation is that catchment works will show up as urban creeks catching up, that signal is not in the data yet.
Water quality is, for the most part, not changing in any way that can be established. One parameter of eleven — water temperature — shows a trend that survives every test, at about half a degree per decade averaged over the year (chapter 6). That single rate is an average over two: the warmer half is warming about three times as fast as the cooler, about a degree per decade against about a quarter, and the whole-year figure is weighted by a sampling calendar that moved rather than by anything about the water. It was offered as a regional climate signal, but the regional climate data do not confirm it: gridded maximum air temperature over the same catchments and the same years rises by only 0.20 °C per decade (-0.17 to 0.56, 28 water years), an interval spanning zero. Conductivity is an uninformative null rather than a reassuring one: +14%, with an interval running from −3% to +34%, which is wide enough to contain a real rise.
And the rare families are being lost. Rare-family richness per sample falls by 21% per decade across all sites, and by 30% per decade on the 45 creeks monitored in twelve or more years (chapter 10, Table 10.4). This is the one negative finding in the report that is demonstrated against the sampling-effort artefact rather than merely adjusted for it: rarefying every sample to a constant 50 individuals — holding the number of animals counted fixed by construction — makes the decline steeper, at 29% per decade (95% CI 8% to 44%, p = 0.008). Sampling effort rose over the record and more animals counted means more rare families detected, so the artefact runs against this finding rather than producing it.
Within the rare set, the sensitive families fall faster still — 48% per decade on the twelve-year panel — but that particular figure is not robust to where the line between “rare” and “not rare” is drawn. Move the cut from 20 samples to 30 and it reverses to +13% (p = 0.39), and the five families that reverse it are recorded almost entirely after 2011, which is the signature of a family entering the laboratory’s working list part-way through the record rather than of a population recovering. The all-rare decline is robust to the same test (−16%, −21%, −12% at cuts of 15, 20 and 30; both sweeps are Table 10.6). The honest form of the conservation argument therefore rests on the rare families as a whole and on the named creeks, not on the −48%. Chapter 10’s Section 10.6 is what we think you might do about it, and R7 is the one-line version.
19.3 Changing the rating system
19.3.1 The decision in front of you
The current system is a reasonable design that has given you a consistent public language for a decade, and nothing here recommends abandoning it. Four factors, a 0-to-5 scale and an average all stay. What changes is how the four factors are measured, how they are scored, how many samples go into a published rating — and how many words the result is reported in.
It splits into three pieces, and you can take the first without committing to the others. Each suggestion is stated in full at Section 19.4; this is the shape of the choice rather than the case for it.
Piece one — reporting. No change to the index at all. Rate from a rolling three-year window rather than from a single sample, and always publish the four factors beside the composite (R1, R6). It is a change to a reporting query, and it is the largest single improvement to the rating’s reliability available to you. It does not get the published word to chapter 16’s 20% criterion, and no achievable number of samples does — which is why R1b asks how many words to publish rather than how many samples to average.
Piece two — the four factors. One afternoon of code. Rarefy the two count factors, replace SIGNAL-SF with the tolerant share on SIGNAL 2, and score each factor once between two anchors rather than twice into integer bands (R2, R3). The case for it is freedom from artefact and not an improvement in discrimination (Section 19.4.2). Its cost arrives with the decision rather than after it: rarefaction cannot score the smallest samples, and those are disproportionately the worst creeks, so R13 — publishing the count of ratings withheld — is a condition of R2 rather than a separate good idea.
Piece three — the anchors. A governance change more than a technical one. Reset the percentile anchors on a rolling ten-year window, generate the table from a versioned script that asserts its boundaries are strictly increasing, and split it by distance from source (R4, R9, R8). One ordering condition attaches, and it is the reason the pieces are numbered rather than listed — do not do the reset without piece two’s scoring rule (R4a).
19.3.2 The three questions you asked, answered as decisions
Should water quality be part of the rating? No — publish it beside the rating. Chapters 6 and 9 look as though they disagree and do not: chapter 6 asks what water quality tells us and answers “quite a lot”; chapter 9 asks whether it separates creeks better than the animals already do and answers “no”. A rating is a compression, and adding an input to a compression only earns its place if it moves creeks relative to one another. Water quality does not — it measures the same urbanisation gradient the animals already respond to, less directly and with more missing data. But diagnosis is not compression, and chapter 6’s parameters are exactly what someone needs in order to act on a bad rating. R5 has the design.
Should the bands be reset on the reference and slightly disturbed sites? Yes, and the fires are not a reason to delay — though not for the reason that first suggests itself. The 2019–20 fires were flagged all through the report as a threat to any re-derivation, because nine of the fifteen delineated reference catchments burnt over more than 90% of their area. The obvious test is to compare bands derived before the fires with bands derived after, and it says the post-fire bands are more demanding, not less — apparently a benchmark that did not move. That test is confounded by which creeks are in each window (Table 14.5): the window before the fires holds 37% burnt-catchment samples against 16% after, and 53% slightly disturbed against 81%. It compares different sets of creeks, not the same creeks before and after.
The test that holds the creeks fixed says the opposite. A within-site difference-in-differences at burnt reference and slightly disturbed sites says they lost 0.47 points of published score (95% CI 0.04 to 0.90) across the fire, relative to unburnt ones, driven by SIGNAL-SF and %EPT. The fires did lower the reference benchmark where the catchment burnt (chapter 14). How much they lowered it depends on how burn extent is measured — taking it from each sample’s own three-year fire history rather than from the 2019–20 season gives a smaller loss, 0.38 points, on an interval that reaches past zero (-0.13 to 0.89) — but the direction does not, and where the periods are cut is not a free choice at all, because every 2020 sample at a burnt site was taken after the fires had started.
Keeping the burnt catchments in the calibration set is still the right call, and now on measured grounds rather than on an absent effect: excluding them from a ten-year window moves 2.7% of rated samples and one site in 77, and it moves them down, because the wet post-fire years lifted the unburnt reference sites by more than the fire depressed the burnt ones. Three conditions attach — keep the window at ten years, publish the excluded band set beside every reset (both R4), and record the window and the burn with the table (R4b) — and one sequencing condition, R4a.
The reset moves 27 of 77 sites down one class. That is a correction to a scale, not a decline in the creeks, and it needs saying that way before it is published, not afterwards.
Should a riparian and geomorphic rapid assessment follow? Yes, as a third panel. You already collect most of one, and the gaps are in consistency rather than in coverage — so a written protocol, fixed photo points and an annual calibration session for the fields already on the form buy more than new fields would (R10, and Table 18.2 for what to add afterwards). Use it as the third panel of a site report — condition (the animals), water quality (trigger exceedance), habitat — and not as a rating factor. A creek that rates Fair with good water quality and poor habitat needs different work from one that rates Fair with good habitat and poor water quality, and a single averaged number cannot tell you which one you are looking at.
19.3.3 And test the next version before you adopt it
You also asked whether a new draft rating system can be tested statistically. It can, on six criteria, and chapter 16 sets them out as a standalone protocol and runs both the current system and the proposed revision through them. The demonstration is a result and not only a worked example: both systems fail reliability against the protocol’s own 20% mark, and the revision fails band behaviour at 10.7% where the incumbent clears 10% at 9.4%. That is two failures, not three.
An earlier version of this chapter said the incumbent also won responsiveness, and that verdict is withdrawn: neither system’s fire response clears zero once it is fitted at the catchment and the year, which is where burn extent varies, and the two are not distinguishable from each other in any case. Tested as the comparison it is — the contrast on the difference of the two standardised scores, over the 924 samples both systems can score — the difference is +0.09 standard deviations (−0.09 to +0.28), an interval that covers zero comfortably. Declaring a winner because one coefficient clears p < 0.05 and the other does not is a comparison of two separately fitted standard errors, not a test of the difference between them (chapter 15, Section 15.4.1). A protocol that our own recommended design does not pass outright is doing its job. That is R12 — and chapter 16 suggests a seventh test no statistic can supply: does the index agree with what your ecologists say when they are standing in the creek?
19.4 What we suggest, in priority order
Seventeen things, ordered by what each buys for what it costs, and indexed in Table 19.1. The identifiers are frozen. R1b, R4a and R4b are conditions on R1 and R4 rather than separate ideas, and nothing has been renumbered, because correspondence has already gone out quoting these numbers — a reader with an old email who looks up “R4a” should find the same thing they were sent.
The strength labels are the report’s own verdicts on the evidence, not on how much we like the idea: evidence compels means the data leave little room to decide otherwise, evidence supports means the case is good and a judgement remains, judgement means the reasoning carries it rather than a measured result, and client request means you asked for it and we have checked it is a good idea.
R1 — Publish the rating from a rolling three-year window, not a single sample. Evidence compels. 38% of same-day replicate pairs land in different rating classes today (95% CI 29% to 47%, over 144 two-sample visits at 59 sites), and it is worse where it matters most — 59% of replicate visits rated Poor disagree about the word, against 18% rated Excellent. Averaging three samples cuts the probability that an independent re-measurement changes the published class from 60% to 41% — a reduction of about 32%. That is a large improvement that still leaves the word unreliable, because chapter 16’s criterion is 20%. It is worth doing whether or not you change the index: it changes no field method, no laboratory step and no arithmetic, and it does not wait on R2, R3 or R4. Effort: none — it is a reporting query.
R1b — Publish three rating classes rather than five, or publish the score and its uncertainty beside the word. Evidence supports. The other half of the same fix. On five classes a three-sample rating changes the word 41% of the time and reaching 20% would take about thirteen samples — 10 to 20 once the uncertainty in chapter 13’s fit is carried — which nobody is going to collect. Neither index repairs this, so it is not a decision that waits on R2 or R3. On chapter 15’s six-way comparison (Table 15.8) the current system changes the published class on three samples 41% of the time and the revision 46% — both more than double the 20% mark. On three classes the three-sample figure is 23% and the five-sample figure 18%. How many words to publish is a communication decision and not ours to take, but it is constrained by this. Effort: none to compute. A decision for you.
R2 — Replace the two count factors with rarefied versions, and SIGNAL-SF with the share of individuals in tolerant families graded on SIGNAL 2. Evidence supports, on freedom from artefact. With the sampling year controlled — the cleaner comparison, because it uses samples from the same site in the same year — the within-site slope of the score on log sample abundance falls from +0.41 to +0.03 standard deviations per natural-log unit of animals counted, an interval of −0.04 to +0.10 that covers zero. Without the year term the same pair is +0.48 and +0.13. Quote them as pairs, never strung into a chain — each form has its own baseline. Discrimination is not the case for it: see Section 19.4.2. Two costs travel with the change and should travel with it in public as well: the revision is slightly worse on independence and on boundary sensitivity (10.7% against the incumbent’s 9.4%, where chapter 16’s mark is 10%), and it withholds a rating from 30 samples — see R13. Responsiveness to fire used to be listed here as a third cost and is not one: neither system’s fire response clears zero once it is fitted at the catchment and the year, which is where burn extent varies, and a test of the difference between them gives +0.09 (−0.09 to +0.28), so the two are not distinguishable (Section 15.4.1). That is one fewer stated cost, not a further argument for the change. Effort: one afternoon of code; no new field or laboratory work.
R3 — Score each factor once on a two-anchor continuous scale instead of twice into integer bands. Evidence supports. It removes the published gaps and the one overlap, doubles the resolution, and gives the 0-to-5 scale a plain reading. The reference comparison adds no ranking information as a second score. Effort: the same afternoon as R2.
R4 — Reset the percentile anchors on a rolling ten-year window of reference and slightly disturbed sites, refreshed every five years, keeping the burnt catchments but publishing the burnt-catchment-excluded band set beside them. Client request, and the current urban bands are more lenient than the distribution they describe. The reset moves 27 of 77 sites down one class and none up. The 2019–20 fires did lower the benchmark at the catchments that burnt, by about 0.47 points (95% CI 0.04 to 0.90) measured on the 2019–20 burn extent — a three-year burn measure gives a smaller loss, 0.38, on an interval reaching past zero (-0.13 to 0.89), so the direction carries more weight than the size (chapter 14). Either way the ten-year window is load-bearing: on that window, excluding the burnt catchments moves 2.7% of rated samples and one site in 77, but on a five-year window a single fire season would move far more. Effort: low, once the table is generated from code — which is R9.
R4a — Do not adopt R4 without R3. Evidence compels. The band reset taken alone makes the reference comparison less informative than it is now: it fixes the pile-up in the urban SIGNAL-SF column, but it makes reference EPT families worse, taking the share of samples scoring exactly 1 from 57% to 72% (Table 14.9). The two belong together. Chapters 14 and 15 state this from their own sides. Effort: none — it is a sequencing condition.
R4b — Record the calibration window, the burn extent and the burn severity alongside the band table. Judgement. So that any published rating can be traced to the benchmark it was scored against, and the next reset can be argued from evidence rather than from assumption. Severity as well as extent is now possible: the 10 m FESM severity rasters for 2019–20 and 2013–14 are in hand. Effort: low, and it is part of R9’s script.
R5 — Publish water quality as a separate trigger-value report card rather than as a rating factor, scoring only alkalinity, nitrate-N and conductivity. Evidence supports. It resolves the apparent conflict between chapters 6 and 9: water quality is diagnostic but does not separate creeks better than the animals already do, and 45% of macroinvertebrate samples have no same-day nitrate-N reading. Score it per parameter and never pooled — an unweighted pool is two thirds probe parameters, so it moves when the probe is replaced. Conductivity carries a condition the other two do not. It is the only one of the three on the probe and its level steps at both probe replacements, so it can be scored only within an instrument era, or against a baseline re-set at each changeover (chapter 6). Phosphate must be reported as detectability, labelled as a property of the test rather than of the creek: the card as originally specified would publish an 89% phosphate pass rate for 2022–2024 of which 80% of the passes are exact zeros. And label the parameter, because its name does not: neither database records whether phosphate is reported as PO4 or as P, the two differ by a factor of three, and every phosphate figure in this report is species-ambiguous in consequence (chapter 7, Section 7.5). Nitrate-N carries its species in its name; phosphate does not. Effort: moderate; the trigger values already exist.
R6 — Report the four factors alongside the composite, always. Evidence supports. The factors do not move together, and the composite’s improvement ends at a changepoint of 2014.4 (95% CI 2010.5 to 2017.25) with no detectable change since. That flat decade is not a quiet one: underneath it, effort-standardised family richness falls -0.85 families per decade over 2010–2024 (-1.65 to -0.05) while the sensitivity component does not fall with it. A composite whose components move in opposite directions and whose average therefore reads flat is the case for this recommendation, not an illustration of it. A composite that hides its components cannot explain itself. Effort: none.
R7 — Flag sensitive rare families, and set up a protection tier separate from the condition rating. Evidence supports. Thirteen rare sensitive families are invisible to all four factors. The specific repair that has been put to us — a 0.25-point bonus on the score for any sample holding one — was fitted in chapter 15’s Section 15.3.6 and rejected there: it changes the published class for about 1% of samples, so it does nothing for the families and spends the score’s plain reading to buy it. A flag is the right instrument and a score is not, because a condition index is built to summarise the common assemblage and short-range endemics are by definition not in it (Harvey 2002). Publish the names beside the rating — “this site supports sensitive rare families: Eustheniidae, Austroperlidae, …” — and report the tier in the habitat panel of R10’s three-part site report rather than anywhere inside the condition rating: chapter 18’s Section 18.3 places it there because the sites that hold these families, the one established Paramelitidae population among them, are not identified by a condition rating and never will be, several of them not being in good condition. Chapter 10’s Section 10.6 is the full version, and it is the one suggestion in this list that is not about the rating at all. Effort: low.
R8 — Split the anchors by position in the drainage network — by distance from source, not by altitude — and settle where to cut before you cut. Judgement: the direction is supported; the cut point is not yet. This suggestion used to name altitude and a 500 m threshold. Altitude was a proxy, and a poor one. On its own it has no relationship with the revised score at all — 0.002 points per 100 m, p = 0.94 across 82 sites — and it becomes visible only once distance from source is held constant, because the sites furthest down their streams are also the lowest ones. Distance from source is the term that holds. A site twice as far down its stream from the source scores 0.26 points better at equal imperviousness (p < 0.001, 82 sites), 0.21 with your own disturbance tiers in the model (p = 0.002), and 0.33 among urban sites alone (p < 0.001, 52 sites) — the subgroup a split table would actually move, and the one where altitude is not distinguishable from zero (0.05 per 100 m, p = 0.167). So there is a real gradient and it is worth anchoring against; a rating that ignores it penalises headwater creeks for being headwater creeks. What is not settled is where to cut. The 500 m was your conductivity and salinity trigger threshold, and it is a threshold on the wrong variable; nothing in this report has fitted an equivalent on distance from source, and one should not be invented by analogy. Choosing it is the first piece of work this suggestion needs, and it should be chosen on the split anchor tables the choice produces — and against what R4’s calibration window leaves on each side of it (Section 19.4.1). The full remedy is chapter 9’s expectation model, which stays deferred for a reason that has nothing to do with whether the effect is real: the catchments are not all delineated. Read the split table as the interim step, not as the answer. Effort: low to split the table once the cut point exists; the cut point itself is an afternoon’s analysis, not a survey.
R9 — Generate the band or anchor table from a script, from a named dataset and a named window, with an assertion that the four boundaries are strictly increasing, and version it. Evidence compels. The published urban bands cannot be reproduced from either database, so they cannot be audited; and a percentile generator without a tie guard produces an empty band on an integer-valued factor whenever the calibration set is small. One band set in chapter 14 does exactly that — the before-the-fires row of Table 14.1, with boundaries of 5, 5, 6 and 8 on EPT families, which makes a score of 2 unreachable. Effort: low, and it stops the next revision costing what this one did.
R10 — Build a riparian and geomorphic rapid assessment as a third report panel, starting with a protocol for the fields already collected. Client request, supported by chapter 9. Not a rating factor: adding shading to the rating reduces its discrimination. The gaps are in consistency, not coverage. What chapter 9 supports is the ordering — habitat ahead of water quality, ahead of imperviousness — and that ordering is now tested against a null and holds. It is also all this suggestion needs. What was withdrawn on 20 August is the stronger form: habitat’s share is not an order of magnitude larger than catchment imperviousness’s, because the noise floor scales with how many variables a block contains and habitat has thirteen against catchment condition’s one. Nothing in the case for a habitat panel rested on that comparison. Effort: moderate, and mostly organisational.
R11 — Design a before-and-after impact study around one of your own works programs. Judgement. The archive holds one designed impact comparison, of three visits. The Leura Falls Creek program is the defensible candidate — one catchment, seven systems with construction windows (Table 17.8), and monitoring either side of them. Nothing else in this list would do as much for the credibility of the rating, and chapter 17 sizes what it can and cannot detect before you commit to it. Effort: high, and worth it.
R12 — Require any future revision to run the six tests and publish the results before it is adopted. Evidence compels. Chapter 16 sets them out as a standalone protocol and demonstrates them on both the current system and the proposed revision, so the comparison is like for like. This suggestion costs nothing and is the only one that protects you from us. Effort: low; the protocol is written.
R13 — Publish the count of withheld ratings, by site, alongside any rarefied rating. Evidence compels. Rarefaction cannot score a sample holding fewer than 20 animals, so R2 withholds a rating from 30 samples. Those are not a random 2.0%: they include 9 of the 31 Very Poor samples in the analysis set (Table 15.4) — the 1,503 rated stream edge samples, which is the narrowest of the several populations a Very Poor count can be taken over. Chapter 1 gives the wider ones and they are not interchangeable. Withholding them silently would make the network look better for a reason that has nothing to do with the creeks. (20 is the rating system’s rarefaction depth and it is deliberately lower than the 50 used for the ecological results earlier in this chapter — a rating has to be issuable for almost every sample, an ecological estimate does not. Chapter 3 reconciles the two.) Effort: none, and it is not optional.
R14 — Call the wetland score a within-Blue-Mountains ranking wherever it is published, and find a second reference wetland. Evidence supports. Everything in this list until now has been about the creeks; the nine wetland sites — six of them points on Glenbrook Lagoon — are rated on a scale that cannot mean what it says, because Ingar Dam is the only reference wetland and one site cannot anchor a percentile band. Relabelling costs nothing and stops the number being read as a condition assessment. A second and ideally a third reference wetland is the one change that would let the wetlands be assessed the way the streams are. Glenbrook Lagoon can now go into the management tables it used to be excluded from, because its six points share one catchment and so carry one imperviousness figure instead of six. Read that figure as the choice of one defensible method rather than as a precise number. The range published beside it is a sweep over two assumptions and not a confidence interval, and it does not contain the three methods’ own disagreement about the same catchment — the land-use estimate sits above the top of it (Section 11.7). That is enough to place the lagoon in a management table; it is not enough to rank it finely against the streams. Chapter 11 has five smaller items besides these two. Effort: low for the relabel; the reference wetlands are a field decision.
19.4.1 Where these pull against each other
Seventeen things written one at a time will interact, and six of the interactions are worth knowing before anything is adopted. None of them is a reason not to act; each is a reason to act in a particular order.
R1 is bought with the network trend, and nobody has priced that until now. Rating from three years of samples makes each site’s word more reliable. At the 70 band-applicable samples a year the record actually carries, doing it by visiting fewer sites more often moves the two things in opposite directions: 70 sites once each gives a standard error of 0.100 on the network mean and 0.543 on one site’s rating; 23 sites three times gives 0.148 and 0.314 (chapter 13, Section 13.8, Table 13.17). If the three-year window is filled by waiting rather than by reallocating the budget, the trade does not arise at all — which is the version we would suggest, and it is why R1 is costed at nothing.
R4a, the one hard ordering condition. Do not reset the anchors without the two-anchor scoring rule. Stated in full at R4a and from both sides in chapters 14 and 15.
R2 changes what R4 is calibrating. The band reset in chapter 14 is derived on the current four factors. Adopt the rarefied factors and the anchors have to be re-derived on the new ones — it is the same script and the same window, but it is not the same table, and running the reset first and the factor change second means doing the reset twice.
R8 makes R9’s tie guard load-bearing. Splitting the anchor table by position in the network halves the calibration set on each side of the cut — wherever the cut ends up falling — and an empty band is exactly what a small calibration set on an integer-valued factor produces. R8 without R9 is how the 5, 5, 6, 8 problem happens again.
R8 halves the calibration set R4 has already sized, and this is the harder of the two. It is not dissolved by R8 moving from altitude to distance from source: the arithmetic is about how many reference sites a split leaves, not about which variable does the splitting. R4’s ten-year window is chosen to be just wide enough. Only nine reference sites enter that window at all; six of them burnt in 2019–20, and R4’s own argument for keeping the burnt catchments in is that dropping them would leave three — a calibration set too small to be stable. R8 cuts the same pool in two from the other direction. Neither suggestion sized the set the two of them produce together, and the split table has not been fitted. So R8’s cut point is constrained twice over: it is untested on the data, and wherever it falls each side of it has to be left with enough reference sites to anchor a percentile. Choose it against whatever window R4 settles on, not independently of it — and if no cut leaves both sides anchorable, that is the finding, and R8 waits on chapter 9’s expectation model rather than being forced through.
R12 is the test R2 and R3 themselves have to pass, and on today’s data they do not pass it outright. The six-test protocol is a suggestion in its own right, so which comes first is a decision and not a formality: adopting R12 ahead of R2 and R3 is a way of deciding to wait for the re-run on 2025–2029 data that chapter 18 asks for, not a way of blocking the change. The run itself, what it says about both systems, and why a protocol our own recommended design does not sweep is doing its job rather than failing, are at Section 19.3 above and in chapter 16 (Section 16.1). Nothing here restates that argument; it is made better where it is made.
One thing that is not a conflict, although it looks like one: R2 is sometimes read as making the index less responsive to catchment fire, while R4b asks you to record burn extent with the band table. R2 has not been shown to change how the index responds to catchment fire: neither system’s fire response clears zero once it is fitted at the catchment and the year, which is where burn extent varies, and the paired test of the difference between the two responses gives +0.09 standard deviations (−0.09 to +0.28), an interval that covers zero (Section 15.4.1). They are different mechanisms in any case — R4b is about being able to reconstruct which benchmark a rating was scored against, not about the index detecting fire.
19.4.2 Why the discrimination result is not the case for R2
It is tempting to sell R2 on discrimination, and we are not going to. The system’s ability to tell a reference creek from an urban one does rise, from 0.76 to 0.85. But:
- it rests on 11 reference sites against 52 urban ones;
- the bootstrap interval on the difference, +0.01 to +0.19, only just clears zero, and a DeLong test on the same pair does not (p = 0.07);
- restricting the comparison to 2020–2024 site means reverses it, 0.78 for the current system against 0.72 for the revision — on an interval of −0.16 to +0.02 that covers zero (DeLong p = 0.22), and it is one evaluation window of nine — 7 of the other eight favour the revision (Section 15.4.1.1);
- and chapter 13 measured how much a single label is worth (Table 13.12): moving 2 sites in 63 between the reference and urban pools the first bullet counts moves this AUC by 0.08, which is the same order as the difference being claimed.
So the tier labels are the measuring instrument here, and they are not accurate enough to referee a contest this close. AUC can establish that a rating works. It cannot pick between two ratings that are a few hundredths apart. R2 stands on freedom from artefact, where the effect is large and the instrument is not in doubt.
19.5 What we could not close
We should be equally plain about what this analysis could not settle, because a conclusions chapter that admits nothing is not being straight with you. Chapter 2 is the full list — Section 2.3 sets out what each gap blocks, and Section 2.9 draws out the three rules the list keeps arriving at. In short:
- How much of each sample the laboratory actually picked is unrecorded, and it is the largest unresolved confound in the report. It is why two of the four rating factors partly measure effort, and why we can adjust for it but not remove it. A pick count on the laboratory sheet closes it going forward, though nothing recovers it retrospectively.
- The detection limits behind the laboratory results are inferred, not confirmed — phosphate reads as an exact zero in about half of all samples, and we have had to guess where the instrument stopped. Your own limits and the test-kit purchase history would settle it, and until they do, every phosphate pass rate in the report is really a detectability rate.
- The probe changes are not separable from the water they measured. Turbidity drops by a factor of eight across two instrument changes and no side-by-side run exists. Two days of overlapping readings at the next change makes the problem never happen again; it does nothing for the twenty-six years already recorded.
- The stormwater asset register is half a register. It has dates for enough assets to make a network-wide evaluation arguable — chapter 17 puts the best detectable effect at about 0.19 score points, with a floor of 0.16 that no amount of extra sampling removes — but it has served catchments for only seven systems and decommissioning dates for none. Those two fields, not more assets, are what would make the works evaluable.
- The published urban bands cannot be reproduced from either database, so we cannot audit the scale the whole rating is measured against. Whoever holds the spreadsheet behind Table 1 of the June 2025 methods document holds the answer.
One thing here has a date on it. Your Planet imagery is licensed through the NSW Imagery Hub and that licence is funded only to the end of FY2026-27; the Nearmap vertical archive over the LGA runs from January 2010 to November 2025 and is available now (chapter 18, Table 18.4). Anything either of them is wanted for — a catchment-scale imperviousness time series is the obvious one — is cheaper to pull while the licence is live than to buy afterwards. It is the only hard deadline anywhere in this report, and it needs a name against it.
None of these is a criticism of the record. Every one of them is a field that nobody had a reason to want until somebody analysed twenty-six years of data at once.
19.6 The data behind this chapter
Each question and request below is set out again in What we need from you, with what it blocks, what an answer is worth and what it would cost you to find, ranked against every other ask in the report.
19.6.1 Questions only you can answer
If the rarefied factors are adopted, will you publish the count of samples that get no rating, by site, alongside the ratings that are published?
Rarefaction cannot score a sample holding fewer animals than the rarefaction depth, so the revised system withholds a rating from 30 samples in the analysis set — and they are not a random 30, because 9 of the 31 Very Poor samples in that same set are among them. That 31 is the Very Poor count over the rated stream edge samples, which is the narrowest of the several populations such a count can be taken over; chapter 1 gives the wider ones and they are not interchangeable. Publishing the surviving ratings without the suppressed count would make the network look better for a reason that has nothing to do with the creeks. We think this is a condition of adopting R2 rather than an optional extra (it is R13), but the commitment has to be yours, because it is your report card.
Refer to it as
dq:withheld-ratings-published.
Which imagery do you want pulled before the NSW Imagery Hub Planet licence lapses at the end of FY2026-27, and who is going to pull it?
This is the only hard deadline anywhere in the report. Planet imagery reaches you through the NSW Imagery Hub and that licence is funded only to the end of FY2026-27; the Nearmap vertical archive over the LGA (18 surveys, January 2010 to November 2025) sits on your separate Nearmap subscription and is available now, and Nearmap AI Packs, Nearmap 3D and Planet Planetary Variables are not. The obvious candidate is a catchment-scale imperviousness time series, which would turn imperviousness from a single modelled snapshot into a covariate that changes over the record — but the decision is what you want, and the answer only has to arrive before the money stops. Chapter 18 Section 18.6 has what is and is not covered.
Refer to it as
dq:imagery-before-licence-lapses.
19.7 A closing note
The Blue Mountains monitoring program has produced a twenty-six-year record that supports a real conclusion: the creeks recovered, substantially, and they have stayed up. What the record cannot tell you is whether they are still moving. The decade since 2014.4 shows no detected change, but the network is not big enough to have detected a slow decline if one had begun — the interval this chapter prints two sections above reaches −0.46 points per decade, which over a decade would take back 52% of the recovery — more than half of it — and would look exactly like what we see. That is why R1 and the monitoring-design questions in chapter 18 matter more than they look. Very few councils in Australia can say the first part with evidence.
The strongest single result in the report is not the recovery itself but its composition — two in five animals in a Blue Mountains creek used to be pollution-tolerant and now fewer than one in five are — and that result got stronger under audit rather than weaker.
None of this is a criticism of the program. It is the ordinary consequence of a long monitoring series being analysed properly for the first time — and most of it is cheap. The first three gaps in Section 19.5 would do more for the next twenty-six years than anything else in this report; between them they cost a few days, and they close the three confounds it has had to work around rather than resolve.
And one thing stands apart from all of it. The condition rating measures what it was asked to measure, and it is silent on the thing this analysis found by accident: the rare and sensitive animals in a handful of named creeks are disappearing, and some of them may exist nowhere else. That is not a rating problem. It is a conservation decision, and it is yours to take.