16  How to test any future rating

Blue Mountains City Council Healthy Waterways — statistical analysis

16.1 A test protocol for any future rating system

Nothing on the master list of data questions comes from this chapter, and that is deliberate. Every other chapter opens with a section of asks; this one has none, because it fits nothing and uses no data of its own — it tells you what to run on yours. The two things its tests do need from you are already on chapter 13’s list: the expert creek rankings the seventh test checks against (dq:expert-creek-rankings), and the sample-versus-visit grain test 2 is judged on (dq:rating-grain-mismatch). Duplicating either here would put one ask on your list twice.

This section answers your question directly and is written to be lifted out. It is a procedure anyone on the team can run against any candidate rating — the revision in chapter 15, a later revision, or a system borrowed from another council — using the data layer that accompanies this report. Each test names what to compute, what a good answer looks like, and what to do if the answer is bad.

The whole procedure needs one input: a vector of candidate scores, one per macroinvertebrate sample, joined to health_score. Call it s.

16.1.1 The scoreboard as it stands

Every test below quotes where the current system and the revision landed, so here is the scoreboard in one place first (Table 16.1). These are chapter 15’s numbers, read from the same fitted object chapter 15 prints in Table 15.8 — not retyped, so the two chapters cannot drift apart. The middle column is the intermediate case chapter 14 builds: the current four factors, scored against percentile bands re-derived from the recent record.

Table 16.1: The two systems, and chapter 14’s intermediate case, against the six tests below. Read from chapter 15’s fitted comparison (Section 15.4.1), over 1,473 edge-habitat stream samples that are not empty and carry a published rating — the same samples for all three systems, the 30 that the revision cannot rate being excluded from every row rather than from one (Section 15.3.4, F-17.44). Discrimination is on 2010–2024 site means over 11 reference and 52 urban sites. Neither system passes test 2. The revision has the better point estimate on discrimination and effort sensitivity, and the incumbent has the better point estimate on reliability, independence and band behaviour. On responsiveness the row shows two separately fitted coefficients and the two systems are NOT distinguishable from each other — see test 5. Every score column fits the three systems separately, so no row of scores is itself a comparison. The last column is the comparison, and it asks the only question Test 1 permits — whether the revision’s difference from the incumbent is separable from zero — from a 400-replicate bootstrap over sites for reliability, class changes, effective dimensionality and the band shift, the stored bootstrap and DeLong tests for the two AUCs and imperviousness, and a paired contrast on the difference of the two standardised scores for effort sensitivity (test 4, and it is also Table 15.9’s last column) and for fire (test 5). The remaining rows are read as comparisons without one. Three of the nine separate, two of them against the revision; read the rest as a direction, not as a margin.
Test Good answer Current Bands reset Revised Revision differs?
1 Discrimination, reference vs urban 0.80 or more 0.76 0.72 0.85 bootstrap yes, DeLong no
1 Discrimination, good vs urban 0.85 or more 0.87 0.85 0.88 no
1 Rank correlation with imperviousness 0.5 or more in size -0.56 -0.54 -0.58 no
2 Reliability (ICC) 0.6 or more 0.49 0.50 0.44 no
2 Class changes on three samples under 20% 41% 41% 46% yes
3 Effective dimensionality, of four factors close to 4 2.59 2.59 2.27 yes
4 Effort sensitivity, year controlled under 0.2 0.41 0.37 0.03 yes
5 Fire response, SD negative -0.18 -0.22 -0.11 no
6 Ratings moved by a 0.1 band shift under 10% 9.4% 12.5% 10.7% no

Four things about Table 16.1 before the tests, because they are easy to miss. No candidate has to win every row — the revision has the worse point estimate on reliability, independence and band behaviour, and the case chapters 15 and 19 make for it rests on discrimination and effort sensitivity rather than on a majority of rows. These are point estimates and not verdicts: chapter 15 adjudicates discrimination and reliability as not established in either direction (Section 15.4.1). The current system already passes three of the seven rows that carry a pass mark, so a candidate that does not beat it on those is a step backwards, not a neutral change. And a row is not a verdictany row, not just the fire one: every score cell in the table is one system’s own coefficient, fitted separately from the other two, so reading a winner off a pair is a comparison nobody ran. That is why the last column and not the three score columns is what the chapter adjudicates on: three of the nine rows put the revision’s difference from the incumbent clear of zero. Effort sensitivity with the year controlled is +0.41 (+0.34 to +0.47) for the current system against +0.03 (-0.04 to +0.10) for the revision — and that row is separated by a test of the difference rather than by eye. The paired contrast over the 1,473 samples both systems can score is -0.38 (-0.44 to -0.32) standard deviations of score per natural-log unit of abundance, -0.47 to -0.29 clustered on the 125 sites, and -0.35 (-0.40 to -0.29) with site fixed effects alone, which excludes zero too. The two slopes say which system sits closer to zero; the contrast says the gap is real. Not all of the separable rows favour the revision. Effort sensitivity, year controlled favours it, and class changes on three samples and effective dimensionality, of four factors separate against it. Two rows have a paired contrast fitted on the difference itself — effort sensitivity (test 4) and fire response (test 5) — and they are the two the chapter is entitled to adjudicate as comparisons. Both were once decided by holding two separately fitted coefficients up against each other, and both pass marks below have been replaced rather than annotated, because this chapter is the protocol: a defective mark carried with a footnote reproduces the error against the next candidate rather than against this one. Read a row as a direction, not as a margin — which is what test 1 to test 6 below tell you to compute, and why the protocol is the deliverable rather than this table.

16.1.2 Test 1 — Discrimination

Compute. Site means of s over the most recent decade. The area under the ROC curve separating disturbance_tier == "reference" from disturbance_tier == "urban", and again separating reference plus slightly disturbed from urban. Use pROC::roc() and report pROC::ci.auc(); where two candidates are being compared on the same sites use pROC::roc.test(..., method = "delong", paired = TRUE), and bootstrap the interval on the difference over sites, not over samples. Also the Spearman correlation of the site means with total imperviousness (ti_primary), with its own site bootstrap.

Good answer. AUC of 0.80 or above for reference against urban; 0.85 or above for good-condition against urban; a rank correlation with imperviousness of at least 0.5 in magnitude. The current system reaches 0.76 / 0.87 / -0.56; the revision 0.85 / 0.88 / -0.58.

If it fails. The candidate is not measuring the condition gradient. Test each factor separately before changing the aggregation: usually one factor is carrying the index and one is diluting it.

Caution, and it is a severe one. Only 11 reference sites have enough recent record to enter this test, against 52 urban ones. On 11 positives the test is very blunt: the 95% interval on the AUC difference between the two systems compared in Table 15.10 runs from +0.01 to +0.19, which is to say the data cannot tell a difference the size of a rounding error from one that would settle the argument outright. Never report the AUC as a bare point estimate, and never adjudicate between two candidates on this test alone. Where the record allows it, derive the index on an early window and test it on a later one. Be careful how you read the result: that split changes two things at once, the window the anchors come from and which samples get judged, so it is not a clean holdout. Chapter 15 sweeps the cut year across 9 splits (Table 15.12); 7 favour the revision, 1 is indistinguishable from zero, and the one published here is the only clear reversal. A statistic that moves that much across 9 reasonable windows on the same data is not measuring a difference between the two indices.

Second caution. The tiers are your own judgement, so this test asks whether the index agrees with the way you have already classified the creeks, not whether it is right — and chapter 13 found errors in that lookup (dq:tier-errors-calibrate-the-test). That is why imperviousness, which is modelled from independent spatial data, is run alongside.

16.1.3 Test 2 — Reliability

Compute. A mixed model s ~ 1 + (1 | sitecode) + (1 | visit_id) + (1 | year) where visit_id is site and date pasted together. The year term is not optional: samples taken in the same year share the weather, the field crew and that year’s laboratory practice, and omitting it charges a network-wide movement to the individual creek. Take the four variance components. Report the intraclass correlation (between-site variance over the total) and the within-site standard deviation, which is everything that is not between sites. Then compute, for the fitted distribution of true site means, the probability that two independent samples fall in different rating classes.

Good answer. ICC of 0.6 or above. A class-change probability below 20% on whatever number of samples the published rating is based on — and if that number is not written down anywhere, that is itself the answer to this test (dq:rating-grain-mismatch).

If it fails. First check whether the bands are too narrow for the noise rather than the index being poor — compare the within-site standard deviation with the 1.00-point class width. Then average over more samples. Then reduce the number of classes (Section 13.7). Both systems tested here fail this test on one sample and on three: the three-sample figures are 41% and 46% against a threshold of 20%, and on five classes reaching 20% would take about 13 samples (19% at that many; 10 to 20 samples once the uncertainty in chapter 13’s fit is carried).

Cross-check. 160 replicate visits exist in the archive, 144 of them with exactly two samples. The observed share falling in different classes should match the modelled probability; if it does not, the variance model is wrong. Quote the observed figure with its interval — over 144 visits at 59 sites it is 38% with a 95% interval of 29% to 47%.

16.1.4 Test 3 — Independence

Compute. Spearman correlations among the factor scores. A principal component analysis of the standardised factor scores; report the share of variance in the first two components and the effective dimensionality, \((\sum\lambda)^2 / \sum\lambda^2\).

Good answer. Effective dimensionality close to the number of factors. No pair of factors correlated above about 0.8. Neither system here gets close: on four factors the current system carries 2.59 dimensions of information and the revision 2.27.

If it fails. Two factors are measuring the same thing and the average is double-weighting it. Either drop one or weight explicitly and say so. Do not leave an implicit weighting in place: the current system weights the EPT signal about twice as heavily as sensitivity, and nothing in the methods document says so.

16.1.5 Test 4 — Freedom from artefact

Compute. Regress the standardised score on log(total_count) with site fixed effects — the within-site slope. Repeat with prop_count_no_sf_grade, and with year restricted to sites monitored throughout, to catch drift in laboratory practice.

When two candidates are being compared, fit the difference. Restrict to the samples both candidates can score, standardise each score on that common set, take the difference sample by sample, and regress that on log(total_count) with the same site — and, in the second specification, year — fixed effects. It is one model and it is the only thing that answers “is this candidate less effort-sensitive than the incumbent?”. Cluster the standard errors on site, or report both currencies: the sites are the unit the fixed effects say the design has. Two separately fitted within-site slopes do not answer it, however far apart their intervals fall — and neither does the ratio of the two point estimates, which has no interval at all.

Good answer. For a candidate on its own: a within-site slope on log abundance below 0.2 standard deviations of score per natural-log unit of abundance. Where a candidate is being weighed against an incumbent, the comparison is passed only if the paired difference between the two within-site slopes excludes zero and the candidate’s own slope is the smaller in magnitude — both halves, because the contrast says whether the gap is real and only the two slopes say which side of zero each system sits on. Report the slope with year fixed effects in and out, as a pair each time, and say which is which; the two specifications have different baselines and are never a chain. The revision reaches +0.13 (+0.07 to +0.20) with site fixed effects and +0.03 (-0.04 to +0.10) with the year controlled, against +0.48 (+0.42 to +0.54) and +0.41 (+0.34 to +0.47) for the current system, and the paired contrast over the 1,473 samples both systems can score is -0.35 (-0.40 to -0.29) and -0.38 (-0.44 to -0.32) respectively (-0.43 to -0.27 and -0.47 to -0.29 clustered on the 125 sites).

This pass mark was wrong until August 2026, in the same way test 5’s was. It gave the absolute threshold and then set the two systems side by side, which left the comparison to be made by eye off two separately fitted slopes. That is not a test of the difference. Two intervals that overlap can still sit either side of a difference that excludes zero; and because the two slopes here are fitted on the same samples they are not independent, so two intervals that fail to overlap are not evidence about the difference either — the sign of the correlation between them decides which way the error runs, and neither the gap between two point estimates nor the ratio of them carries an interval. Chapter 15 carried the identical defect on this criterion and published the contrast in August 2026 (Table 15.9). The mark above replaces it rather than annotating it, for the reason test 5’s was replaced: this chapter is the protocol, so a mark left standing with a caution attached is applied to the next candidate, not to this one. No verdict moves. The contrast excludes zero in both specifications and in both currencies, and it agrees with the two slopes about which system sits closer to zero, so freedom from artefact remains the criterion the revision wins and the one the book’s case for it rests on (Section 15.4.1).

If it fails. Rarefy the count-based factors, or fix the pick count in the field and laboratory protocol. Rarefaction is the cheaper fix and works retrospectively on the whole archive; a protocol change only helps from the day it starts.

16.1.6 Test 5 — Responsiveness

Compute. Within-site change at sites with a documented impact, with a year random effect, for the same reason as Test 2 and with more force: impacts are concentrated in particular years, so without a year term the weather of those years is confounded with the impact. Two impacts are available: catchment fire (site_fire_bug$prop_burnt_3y — the _bug object, not the stacked site_fire_samples, which mixes both databases) and the Hazelbrook bifenthrin control-impact pair (87GHZ and 88GHZ). Also the power to detect a trend: fit the site-level trend model of chapter 5 and compare the standard error on the slope between candidates.

When two candidates are being compared, fit the difference. Restrict to the samples both candidates can score, take the difference of the two standardised scores sample by sample, and fit the impact term on that, with the same site, visit and year terms. It is one model and it is the only thing that answers “does this candidate respond better than the incumbent?”. Two separately fitted impact coefficients do not answer it, however far apart their p-values fall.

Good answer. The right sign, on an interval that excludes zero. Where a candidate is being weighed against an incumbent, the comparison is passed only if the paired difference between the two responses excludes zero, in the candidate’s favour.

This pass mark was wrong until August 2026, not merely wrongly worded. It read “the right sign, and a smaller standard error than the incumbent”. A standard error measures precision, not response: a candidate that does not move at all can have the smaller one, and a candidate can win it while responding less. Applied to the two systems below it awarded the criterion to the incumbent because the incumbent’s coefficient cleared p < 0.05 and the revision’s did not — the difference-between-significant-and-not-significant error. The mark above replaces it, and the verdict it produced is withdrawn (chapter 15, Section 15.4.1). Keep the standard errors in the report; they belong in the power half of this test, which is a question about how much data a candidate needs, not about how much the creek changed.

If it fails, or cannot be run. Say so. On the fire test neither response clears zero once it is fitted at the catchment and the year, which is where burn extent varies: -0.18 (-0.39 to +0.02) standard deviations for the incumbent against -0.11 (-0.32 to +0.10) for the revision, both fitted over the 767 site-years the 924 fire-set samples reduce to, because burn extent is a property of the catchment and the year rather than of the sample and only 14 of the 83 creeks in the set ever burn. A calibrated site permutation over those creeks gives p = 0.093 and p = 0.264. That is not the incumbent winning. The paired difference, over the 924 samples both systems can score (64 of them burnt, at 83 sites), is +0.09 standard deviations (-0.09 to +0.28) — an interval that covers zero comfortably, so neither system is shown to respond to catchment fire better than the other. “The revision responds better” is not supportable either; what is not supportable is a verdict in either direction. All of it rests on 64 post-fire samples out of 924, so this is the criterion most in need of the monitoring design below.

The most valuable single thing you could do for the credibility of the rating is a before-and-after monitoring design around your own works — a stormwater treatment installation, a riparian planting, a creek naturalisation — with control sites and at least three years of pre-works data. Chapter 17 works out what that would take and what it would be able to detect (Table 17.6).

16.1.7 Test 6 — Band behaviour

Compute. Check that the band table is a partition: for every factor, that the upper bound of each band equals the lower bound of the next and that the open/closed flags do not overlap. Assert that the four percentile boundaries are strictly increasingstopifnot(all(diff(q) > 0)) — because on an integer-valued factor two adjacent percentiles can be equal and a tie makes a band empty (Section 14.3.2). Tabulate the share of samples in each band. Compute the share of samples within one quarter of a class boundary. Re-run the whole rating with the class boundaries shifted by 0.1 in each direction and count how many published ratings change.

Good answer. A clean partition, with strictly increasing boundaries. No band holding more than about 40% of samples. Fewer than 10% of published ratings changing under a 0.1 shift of the boundaries — the current system moves 9.4% and the revision 10.7%, so the current system clears the 10% mark and the revision does not, and both are close enough to it that a small refit could move either across.

If it fails. Gaps, overlaps and empty bands are always a bug and should be fixed in code — widen the calibration window until the boundaries separate, or collapse the tied scores into one band and say so in the published table. Piling up in one band means the bands were drawn from a different distribution than the one they are applied to (Section 13.3.1) — but check the direction first, because a pile-up in a reference comparison applied to an urban network is the design working. Excessive sensitivity to boundary placement means the classes are finer than the data supports; Section 13.7 works out how many the data here can carry, and the answer is three, not five.

16.1.8 A seventh test worth adding: does it agree with the ecologists?

None of the above tests whether the rating agrees with expert judgement about particular creeks. That is worth testing and it is easy: ask three experienced staff to rank twenty creeks they know, independently, before seeing the scores, and compare. A rank correlation below about 0.7 with the consensus is worth investigating — it will usually turn out that the experts are using information the index does not have (channel condition, weeds, flow permanence), which is useful to know. The ask is dq:expert-creek-rankings, and it costs an afternoon.

16.1.9 Running order

Run tests 3 and 4 first, on individual factors, before assembling anything. Run 1 and 2 on the assembled index. Run 6 last, once the bands are set. Test 5 is opportunistic and will usually be inconclusive until there is an impact design to run it against. The two paired halves — test 4’s and test 5’s — come after assembly rather than with the rest of their tests, because a difference of two standardised scores needs both indices built and both standardised on the samples they can both score.