2.4 What we are asking for
130 asks in all, from every chapter of the report.
| Group | Questions | Data or documents | Total |
|---|---|---|---|
| For you | 69 | 38 | 107 |
| Recording changes | 2 | 12 | 14 |
| Not for you | 3 | 4 | 7 |
| Purchasing | 2 | 0 | 2 |
A further 4 items on our internal register turned out to be already in hand, or superseded, and have been dropped rather than asked for twice.
2.4.1 The 10 that matter most
| # | Ask | Chapter | Value | Costs you |
|---|---|---|---|---|
| 1 | How many animals was the lab asked to pick, and did it change? | 3 | transformative | minutes |
| 2 | Did the old and new probes ever run on the same day? | 3 | transformative | minutes |
| 3 | Does a zero mean “none there” or “too little to see”? | 3 | transformative | minutes |
| 4 | Which kit was used, and what floor does it state? | 3 | transformative | an afternoon |
| 5 | Where is the Leura Falls monitoring data? | 17 | transformative | an afternoon |
| 6 | What can the phosphate kit detect? | 7 | transformative | an afternoon |
| 7 | Any training note on how to estimate cover, 2010-2016 | 3 | transformative | an afternoon |
| 8 | Is there another reference wetland we could add? | 11 | transformative | an afternoon |
| 9 | One season of sampling beside the state program | 12 | transformative | a real search |
| 10 | The complete drainage network, including what you do not own | 4 | transformative | a real search |
Ranked by what the answer is worth against what it costs you to find. Every one of them is set out in full below.
2.5 The list, by topic
2.5.1 Macroinvertebrates — laboratory
1. How many animals was the laboratory asked to pick out of each sample, was any subsampling used, and on what dates did either change? Even a dated instruction or an email thread would settle it.
The median number of animals found per edge stream sample went from 92 over 1998-2009 to 167 from 2014 on, a factor of 1.39 per decade within a site. That is either creeks getting healthier or somebody picking harder, and nothing in the archive distinguishes them. Adjusting for it removes about a third of the headline improvement in waterway health — the single most important number in the report. An entire chapter’s worth of statistical machinery exists to work around this one missing document, and it throws away a fifth of the samples doing it. If the answer is “there was no protocol”, that is still the answer, and it says start recording it today.
Blocks: The largest confound in the report; the rarefaction apparatus that discards 21.6% of samples. Value: transformative. Costs you: minutes. Refer to it as
dq:pick-count-protocol.
2. On the 160 visits with more than one macroinvertebrate sample, why was the second one taken — a separate collection from the creek, one collection split in two, or the lab checking itself? Registration books or lab submission forms would show it.
The most quoted number in the whole report is that two samples from the same creek on the same day land in different published rating classes 38% of the time, over the 144 two-sample visits — an observed count, rising to 41% of all 160 visits if the 16 with three or more samples are included. That 41% is not the modelled 41% the executive summary quotes for a three-year rating; the two are unrelated quantities that happen to round to the same two digits. What this one means depends entirely on which of those three it is. If they are split subsamples, 38% is a lower bound. If they are lab duplicates, the fix is laboratory QC, not more field visits. You are being asked to spend money on the strength of a number whose meaning is currently unknown, and this decides which budget line it comes out of.
Blocks: Which budget line the sampling recommendation points at; the variance decomposition behind it. Value: transformative. Costs you: a real search. Refer to it as
dq:replicate-purpose.
3. Chironomids vanish from the records in 2008 and appear in only 5 of 36 samples in 2009. Was that a lost data-entry batch, or did somebody change what the lab was asked to identify?
Chironomids average 14% of the individuals in a sample and one to three rows of the family count, so both the family count and %EPT are distorted across those years. The choice is between excluding two years and modelling the change, and we cannot make it without knowing which happened. Somebody may simply remember.
Blocks: Whether 2008-09 is excluded or modelled; any family count or %EPT series crossing 2008-09. Value: high. Costs you: minutes. Refer to it as
dq:chironomid-absence-2008.
4. We found three quirks in the way the family count comes out of CalcsStep1. Which of them are deliberate, and which would you want fixed in the next version?
One, the five chironomid subfamilies are relabelled “Chironomidae” after the SELECT DISTINCT, so a sample with Chironominae and Orthocladiinae counts Chironomidae twice — your own Lapstone Creek worked example counts 15 families including Chironomidae twice. Two, where a taxon is entered twice for one sample with different counts both rows survive and it counts twice, but if the two counts happen to be identical they collapse to one. Three, the family count excludes Collembola, Ostracoda, Copepoda and Cladocera but the abundance and %EPT denominators include them. We have reproduced all three exactly and changed nothing; whether they should change is your call, not ours.
Blocks: Whether the rating review should recommend the strict family count, and whether the percentile bands need re-deriving. Value: high. Costs you: minutes. Refer to it as
dq:family-count-defects.
5. When a sample is subsampled in the laboratory, is what reaches the tray close to a random draw from the sweep — or does anything about the sorting make big, obvious or fast-moving animals more likely to be picked?
The two rarefied factors assume the animals counted are a random draw from the animals present. That assumption is doing real work: if larger or more conspicuous taxa are picked preferentially, rarefaction standardises the count but not the bias, and it would show up as a stable-looking factor that is quietly wrong in the same direction every time. We cannot test this from the database — it needs someone who has done the sorting. Even an informal answer (“we pick until the tray looks done”, “we grid the tray and do squares at random”) would tell us how much weight the assumption can take.
Blocks: How firmly the effort-standardisation case can be made, which is the whole case for the revision. Value: high. Costs you: minutes. Refer to it as
dq:rarefaction-random-draw.
6. When did each family name enter or leave the laboratory’s reference list? A dated version of the list would be ideal — failing that, does anyone remember Antipodoeciidae turning up?
The report’s headline conservation finding is that twelve rare families disappeared before 2011, and three of those have already turned out to be nomenclature rather than extinction. The 48% decline in sensitive rare families is not robust to where the rarity threshold is drawn, and the families that reverse it look like list changes: Antipodoeciidae has 24 of its 26 records after 2011 and none before 2006. This is the last open attack on the most consequential conservation claim in the report, and the reference list is what closes it.
Blocks: The rare-family decline, and whether the dynamic occupancy model would treat a naming change as a colonisation event. Value: high. Costs you: an afternoon. Refer to it as
dq:dated-family-list.
7. Are the field sheets still around for sample codes 1579 to 1587 — nine consecutive codes, all 2018, at nine different creeks, every one of them recorded as holding no animals at all?
Nine consecutive sample codes at nine different creeks all coming back empty is either a remarkable autumn or a lost laboratory batch, and nothing in either database tells us which. They are nine of the twelve empty samples in the entire record. Because they are empty they drop out of every rating chapter’s analysis set before any scoring happens, so they are invisible in the results rather than visible as bad news — which is exactly the wrong way round if the creeks really were that dead. The sites are 36GSP, 48NGK, 50NEP, 52NGKR, 38.2NVH, 47NBX, 41NWL, 46NEH and 44NYK, and the dates run 12 April to 24 May 2018.
Blocks: Whether nine creeks genuinely held nothing in autumn 2018, and how many Very Poor samples the record should contain. Value: high. Costs you: an afternoon. Refer to it as
dq:unrateable-field-sheets.
8. Are the laboratory identification sheets, or a LabIDOfficer roster, still around for 2000-01 and 2008-09?
This is the documentary version of the chironomid question. The bug database only records a lab ID officer from 2019 onwards, so for the two suspicious periods there is no record of who identified what or to what level.
Blocks: Whether two years of the record can be included. Value: high. Costs you: a real search. Refer to it as
dq:lab-id-roster-2000s.
9. Could someone re-export CalcsStep4_EPT from Access with the %EPT column populated? It was a calculated column and it came across empty in all 2,062 rows.
It is the one published quantity we cannot check against the whole series — we match your family counts, abundances and EPT counts on every sample, but %EPT only against the three worked examples. A five-minute export turns the last unchecked figure into a checked one.
Blocks: A full-series reconciliation of %EPT, which is one of the four factors the rating is built from. Value: moderate. Costs you: minutes. Refer to it as
dq:ept-percent-column-export.
10. Two sensitive families run against the general recovery — the stonefly Notonemouridae and the damselfly Synlestidae are both becoming less widespread. Does that match what your field staff see?
Everything else in this chapter says sensitive families are gaining ground, and these two are the exceptions that survive a fully corrected model. They may be a genuine local loss worth acting on, or they may be an identification habit. Field recollection would tell us which is worth chasing before we spend any more analysis on it.
Blocks: Whether either family is worth a targeted look. Value: moderate. Costs you: minutes. Refer to it as
dq:sensitive-family-decliners.
11. Were terrestrial and semi-terrestrial animals — Oniscidae and Talitridae, 21 samples each — always recorded when they turned up, or did the instruction change at some point?
Both families appear to decline over the record. Whether that is an animal or a convention we cannot tell from the data, and one long-serving officer’s recollection would settle it in a minute.
Blocks: Whether the decline in two families is real or a recording change. Value: moderate. Costs you: minutes. Refer to it as
dq:terrestrial-taxa-convention.
12. Some samples carry the same taxon twice with two different counts. Is that two picks of the same sample, a correction that was never deleted, or a typing slip — and would the lab sheets tell us?
We sum them, which is the only choice that does not throw a record away, and the family count is left exactly as your query produces it. But the duplicates are the reason a taxon sometimes counts twice and sometimes once, so knowing what they are would settle the question above as well.
Blocks: Nothing analytically — it would settle the family-count question. Value: moderate. Costs you: an afternoon. Refer to it as
dq:duplicate-taxon-rows.
13. Could the Waterfall Creek Paramelitidae get a specialist identification? And is the field sheet for sample 590 (2010, 100 individuals) still around to confirm that count?
This is the named high-conservation-value population in the report and a protection recommendation rests on it. Unlike most of this list it is new expert work rather than a filing-cabinet search, so it costs real money — but it is a small amount of money against the weight the finding carries.
Blocks: The protection recommendation for Waterfall Creek. Value: moderate. Costs you: a real search. Refer to it as
dq:paramelitidae-verification.
2.5.2 Water quality — laboratory
14. When the lab wrote a plain 0 for phosphate, nitrate-N or faecal coliforms, did it mean “none detected” or “below what the test can see”? And is the 5 written for coliforms from 2020 the same thing?
More than half the phosphate readings in twenty-three years — 53.8% — sit below what the method could see, nearly all of them written down as an exact zero, and a third of the coliform ones. If they mean “below the detection limit”, then the cheerful finding that phosphate improved is mostly a story about test kits, and a published range starting at 0.00 has been silently scoring every non-detect as a pass. If they are real zeros, the finding stands. One person who worked the bench can answer this in a sentence, and nothing else on this list has that ratio.
Blocks: Whether 53.8% of phosphate readings are censored or real; the phosphate exceedance improvement, which currently reads as a change in kit detectability. Value: transformative. Costs you: minutes. Refer to it as
dq:zeros-are-nondetects.
15. Which test kit or method did you use for phosphate and nitrate-N, and what does its packaging or instruction sheet say it can detect? A purchase order, a stores ledger entry or a photograph of a kit box would all do it.
We have had to guess at this, and the guess was wrong once already, which is why it is near the top of the list. Nothing in either database records what the method could detect — so the limit was taken from the smallest value the data held, 0.005 ppm, and that value turned out to be one our own arithmetic produced. Seven phosphate and three nitrate samples between 22 February and 27 April 2012 are each the average of a laboratory duplicate reading 0.00 and 0.01. The raw table holds nothing below 0.010 in any year and every reading sits on a 0.01 grid, so we now use 0.010 — and making that correction moved those ten samples from “measured” to “not detected”. A number that decides whether a reading is a measurement or an absence should not be coming out of our averaging. Two things suggest where to look. tblSamples.LabOfficer names Council staff rather than a laboratory, and the coliform notes describe 1 mL and 10 mL aliquots plated in-house — so this looks like a bench or field test kit, and a 0.01 ppm grid unbroken across twenty-two years is what a colorimetric kit produces. A commercial laboratory would report continuous values against a stated limit of reporting, which is exactly the document we are missing. And the grid only tells us the resolution the results were written at, not the floor of the kit that produced them: the only two field notes that state a floor disagree, “Phos <0.1, Nitrate <0.1” in September 2008 and “Phos<0.01, Nitrate <0.01” five months later, both recorded beside a stored zero. If the kit or its range changed between those two dates, that would be worth knowing too. Coliforms are separately documented at 10 CFU/100 mL in the field notes and need no further work. The same kit box would settle one more thing, asked separately as dq:phosphate-units: whether the phosphate result is written as PO4 or as P. Nothing in either database, the field sheets or the methods document says which, the two differ by a factor of three, and every phosphate concentration in ppm in this report is ambiguous by that factor until it is answered. Your nitrogen field says which it is — it is called Nitrate-Nitrogen — and the phosphate field does not.
Blocks: Fitting a censored-data model; restoring phosphate as a report-card component; whether 53.8% of phosphate readings are censored or real. Value: transformative. Costs you: an afternoon. Refer to it as
dq:lab-detection-limits.
16. Can you find anything that names the phosphate test kit or reagent used between 2002 and 2024, and states what it can detect — a purchase order, an invoice, a stores ledger entry, even a photograph of a kit box? One dated point would help; two or three would settle it.
This is the same ask as chapter 3’s lab-detection-limits, narrowed to the one parameter it would unblock fastest — and it is not quite as small as it once looked. The floor did not move during the record, so there is one limit to confirm rather than a dated history of two kits: the smallest positive phosphate reading is 0.010 ppm in every year including 2012, and the seven readings that sit at 0.005 are averages of two replicates reading 0.00 and 0.01 rather than values the kit ever reported. But that only shows us the grid the results were written on, which is not the same thing as what the kit could see — a kit with a 0.1 ppm floor writing onto a 0.01 grid would leave exactly this trace, and one 2008 field note reading “Phos <0.1” suggests it might have. With an answer, phosphate becomes the strongest candidate in the set; without it, the 89% of the 2022-2024 report-card window’s phosphate measurements that sit inside the desirable range is mostly a statement about the kit.
Blocks: Reporting phosphate as anything other than detectability; fitting a censored-data model to the 54% of readings that are non-detects. Value: transformative. Costs you: an afternoon. Refer to it as
dq:phosphate-kit-history.
17. In 2019, nitrate, ammonium and turbidity all show blocks of exact zeros at the same time. Was a blank cell being entered as a zero that season?
Turbidity is a probe reading with no detection limit at all, so a zero there cannot mean “below the limit” — which is what makes the three parameters moving together suspicious. If it is a data-entry convention, a whole season of three parameters needs flagging as not-measured rather than measured-as- zero, and that distinction is unrecoverable after the fact.
Blocks: Separating “not measured” from “below limit” for a whole season. Value: high. Costs you: minutes. Refer to it as
dq:blank-vs-nondetect-2019.
18. When phosphate is written down in ppm, is that phosphate as PO4 or as phosphorus? The kit’s instruction sheet, a results header or whoever set up the spreadsheet would settle it.
The field is AvailPhosphate and the unit is ppm, and nothing in either database, in the field sheets or in the methods document says which of the two conventions the number follows. They differ by a factor of 3.07, so a reading of 0.03 ppm is either 0.03 ppm of phosphorus or about 0.01 ppm of phosphorus depending on which was meant. Your nitrogen field says which it is — it is called Nitrate-Nitrogen — and the phosphate field does not. Your own 2024 desirable ranges are safe either way, because they were derived from percentiles of these same readings, so both sides of a pass or a fail carry whatever convention the kit used and it cancels. What is not safe is any number that leaves the building.
Blocks: Any comparison of these readings to a guideline or to another program’s results, including chapter 7’s check against an independent laboratory’s total phosphorus limit, which is right or wrong by a factor of three depending on the answer. Value: high. Costs you: minutes. Refer to it as
dq:phosphate-units.
19. For phosphate, turbidity and faecal coliforms the desirable range starts at zero, so a reading the test could not see is scored as a pass. Are you happy for us to print the non-detect rate beside those three wherever the pass rate appears?
It is not a small correction. 70% of phosphate’s “within the desirable range” flags in the analysis set are stored as literal zeros — 71% if you also count the readings below the inferred 0.01 ppm limit, and the chapter keeps those two figures apart because only one of them needs a limit nobody has confirmed. Not one of the 247 “above” flags is a zero. For faecal coliforms, whose detection limit of 10 CFU/100 mL is the one that is actually confirmed, 73% of the passes are below that limit. Printing the pass rate on its own reports the test kit. Printing both numbers is honest and costs one extra column.
Blocks: Any published pass rate for phosphate, turbidity or faecal coliforms. Value: high. Costs you: minutes. Refer to it as
dq:zero-start-ranges.
20. Did the nutrient test kits change over 2017-2019? Anything at all — a supplier change, a new reagent, a different reading method?
The share of phosphate results below the detection limit runs 75-89% over 2013-2016, falls to 17-34% over 2017-2019, and returns to 61-90% from 2020. Creeks do not do that. Something in the method is switching back and forth, and until we know what, none of those years can be read as a nutrient signal.
Blocks: Reading anything from the phosphate record over 2017-2019. Value: high. Costs you: an afternoon. Refer to it as
dq:nutrient-2017-2019-change.
21. Something changed at the bench over the summer of 2004-05 and faecal coliform counts fell about a hundredfold. Is there a method log, an SOP from 2003-2006, an analysis report or an invoice that would say what?
The step is identical at every site, including the unsewered reference catchments, which rules out sewer works and points squarely at the laboratory. Field sheets for the summers of 2004 and 2005 would help too. As it stands, twenty-four years of coliform data cannot be joined into one series across the break.
Blocks: A single coliform series spanning the whole record. Value: high. Costs you: a real search. Refer to it as
dq:coliform-2004-break.
22. How was a below-detection alkalinity value recorded, and did the convention change? A literal 1 appears 38 times in 1998-99 and never once in the other twenty-five years.
Three counts describe this and they count different things, so here they are together. Thirty-eight is raw measurement rows in tblWaterQuality holding an alkalinity of exactly 1, every one of them in 1998-99. Nineteen survive into the analysis set, once replicates are averaged and the seven copy-down placeholders the data layer corrects are taken out. Four readings, none of them a 1, are recorded as exact zero, and those four are the alkalinity non-detects the censoring census counts — the 1s are not flagged as censored, because nothing in the data says they are, and that is precisely the question. No test on this data can separate them from genuine censoring; the answer has to come from a person. Small, but it is one of the few things on this list that is literally unanswerable without you.
Blocks: Whether 19 alkalinity readings are measurements or non-detects. Value: moderate. Costs you: minutes. Refer to it as
dq:alkalinity-1s.
23. How were bounded and estimated plate counts recorded? Three of the largest values in the record went in as “>10,000”, “TNTC” and a conservative estimate entered as though it were measured.
23,000 is the fourth-largest post-break coliform value and it is not a measurement — the field note estimated it “conservatively as this is the highest recorded”, and it was not: the same site had already read 29,000 thirteen months earlier, and 182,000 there in 2017 and 88,500 at 86BKT in 2022 are both larger again. Any coliform maximum or mean published from this record leans on those three numbers, so the convention matters more than the count of affected rows suggests.
Blocks: Any published coliform maximum or mean. Value: moderate. Costs you: minutes. Refer to it as
dq:coliform-bound-convention.
24. Would you like us to switch replicate coliform averaging from an arithmetic to a geometric mean, as ANZECC/ARMCANZ specifies?
It affects sixteen sample codes, so it is small. But the current choice is indefensible if anyone asks, and changing it moves numbers you have already published — which is why it should be your call rather than ours.
Blocks: Sixteen sample codes, and the defensibility of the published means. Value: low. Costs you: minutes. Refer to it as
dq:coliform-mean-decision.
25. Four site/date pairs carry probe readings identical to seven significant figures (two in 2021, one in 2023, one in 2024). Is that one visit entered twice, or two visits?
The water quality table has 2,688 rows against 2,684 distinct visits, so anything counting samples currently weights these four double. Trivially small and trivially cheap, but it should be right.
Blocks: Water quality sample counts. Value: low. Costs you: minutes. Refer to it as
dq:duplicate-visits.
26. Was pH in 2003 recorded from a different sheet, or by somebody rounding? 94% of that season’s readings sit on a half-unit grid, against 12% in 2004 and single figures in most other years.
2008 and 2002 show weaker versions of the same thing. Nothing published rests on it, but 2003 cannot contribute to a pH trend at the resolution the rest of the record supports, and it would be good to know why.
Blocks: 2003’s contribution to any pH trend. Value: low. Costs you: minutes. Refer to it as
dq:ph-2003-coarse.
2.5.3 Water quality — probe
27. Did two probes ever go out on the same site on the same day — even a handful of samples, even informally during a changeover?
Turbidity fell about eightyfold across the two instrument changes, at 54 of the 55 stream sites measured both before and after: the site median goes from 4.73 NTU over 2012-16 to 0.058 NTU over 2020-24. Creeks do not do that; sensors do. Dissolved oxygen steps up 0.79 mg/L at the first change and 0.66 at the second, which is more than its whole apparent improvement over the record. With even a few days where both instruments were used we can stitch the three eras into one 27-year record. Without it, turbidity and dissolved oxygen have simply been abandoned. It is a yes/no question and the answer is worth two whole parameters.
Blocks: Any turbidity or dissolved oxygen trend crossing the 2017 or 2020 instrument changes. Value: transformative. Costs you: minutes. Refer to it as
dq:side-by-side-runs.
28. The share of turbidity readings above the 7.08 NTU trigger falls by a factor of three across the three probes, and dissolved oxygen moves nearly seven-fold the other way. Would you rather we left both out of the report card, or published them with the instrument era printed beside them?
Each probe reads systematically lower than the last, and a probe that reads lower crosses a fixed threshold less often. Chapter 3’s verdict table (Table 3.1) records the turbidity change as not separable from the method. Conductivity is the counter-example that makes the rule usable — it moves 1.5-fold across the same three instruments, which is what a stable measurement looks like.
Blocks: Any published turbidity or dissolved oxygen exceedance rate crossing the 2017 or 2020 probe changes. Value: high. Costs you: minutes. Refer to it as
dq:exceedance-tracks-the-probe.
29. Do you have the manufacturer’s turbidity sensor specification and stated resolution for the Hydrolab Quanta, the Aquaread AP-2000 and the In-Situ Aqua TROLL 500? A data sheet or a manual would do.
We have had to invent a 0.1 NTU detection floor with no documentary basis whatsoever, and three headline percentages in the water quality chapter rest on it. A published spec sheet replaces our invention with a fact.
Blocks: Three published percentages that currently rest on an invented floor. Value: high. Costs you: minutes. Refer to it as
dq:probe-specs.
30. On the Hydrolab and Aquaread, was salinity in PSU computed from electrical conductivity by a fixed factor — and if so, which factor?
Pre-2020 salinity is recorded on a 0.01 PSU grid, and the published desirable range is exactly one quantisation step wide, so no pre-2020 salinity rate or trend means anything at the moment. If PSU was derived from EC by a known factor we can reconstruct it at full precision from the EC column and get the whole parameter back. One number does it.
Blocks: Any use of pre-2020 salinity. Value: high. Costs you: minutes. Refer to it as
dq:salinity-psu-derivation.
31. Is there a working spreadsheet or method note behind the 2024 desirable ranges — which sites, which years, and which instrument each reading came off? We have reconstructed them as 20th–80th percentiles of reference and slightly disturbed sites, 2006–2023, and we would like to check that.
Everything this chapter prints is a comparison against those ranges, so their derivation decides what the numbers mean. Two things we would want to see in it. The dissolved oxygen and turbidity ranges pool the Hydrolab and Aquaread eras, so the published range for those two describes a mixture of two instruments and every future pass rate is measured against the mixture. And the site list is 32 stream sites and two wetlands, so the published ranges are a stream distribution with a token of wetland in it rather than a wetland benchmark — which is why wetland readings are excluded here (chapter 11).
Blocks: Knowing whether a pass rate is measured against a stable benchmark, and whether the ranges need re-deriving after the 2017 and 2020 probe changes. Value: high. Costs you: a real search. Refer to it as
dq:desirable-range-derivation.
32. Which physical instrument do the four device labels and two serial numbers in the database correspond to — and what was used on the 135 Aqua TROLL-era turbidity readings that carry no device identifier?
Housekeeping, but it is what lets an instrument step be pinned to an instrument rather than to a date. Somebody who was there will know this straight away.
Blocks: Attributing instrument steps to instruments. Value: moderate. Costs you: minutes. Refer to it as
dq:probe-device-labels.
33. Are there calibration records for the probes — particularly for the two Aqua TROLL units? They may well be on paper.
The two Aqua TROLLs disagree by 40 percentage points on how often they return a sub-floor turbidity value. Without logs that cannot be attributed to either instrument, and it narrows rather than closes the turbidity problem even if we get them.
Blocks: Attributing the Aqua TROLL disagreement to one unit or the other. Value: moderate. Costs you: a real search. Refer to it as
dq:probe-calibration-logs.
2.5.4 Field protocol and observers
34. Was there ever a training note, induction, worked example or run sheet on how to estimate cover — anywhere between about 2010 and 2016? A dated version of the site description form would also do it.
Recorded riparian shading rises 11.7 percentage points a decade at sites that have not changed, while moss, detritus and trailing vegetation drift down. Aerial imagery has now shown officers are right about which reaches are shady and that their year-to-year movements at a fixed reach do not track canopy — so the rise is in the recording, not the trees. What imagery cannot recover, at any budget, is why. A training record from the mid-2010s would. This matters because the report’s main practical suggestion is “shade your creeks”, and it rests on that variable.
Blocks: Whether the report’s main on-ground suggestion holds; the reading of ten of thirteen habitat variables. Value: transformative. Costs you: an afternoon. Refer to it as
dq:training-record-2010-2016.
35. Would you be willing to run one season’s macroinvertebrate sampling to the AUSRIVAS protocol alongside your own, at the reference sites — and could DCCEEW be asked to sample the same reaches in the same season?
The whole archive holds exactly one visit pair that is genuinely comparable to one of yours: same habitat, same austral season, twelve days apart. One co-located field round at the twelve archive sites in the proximity set would produce twelve same-day pairs — more than the archive has accumulated in thirty years — and run at the eleven reference sites instead it would give the first external assessment of ten of them. The protocols are already compatible: both sample edge and riffle, both identify to family, both use the same SIGNAL grade tables, and they pick to a similar effort (median edge richness 15 against 14). Nothing needs to be developed first. This is the single most concrete data request in the report.
Blocks: Any external validation of the rating, and any defensible statement about the reference set as it stands after the 2019–20 fires. Value: transformative. Costs you: a real search. Refer to it as
dq:ausrivas-colocated-round.
36. The Macro round has drifted from early March to late May over the record. Can it be pinned to a fixed window from now on — and does anything (staff, budget cycle, weather, contractor availability) make one window easier to hold than another?
The median sampling date moved from 2 March over 1998-2005 to 26 May over 2020-2025, about three months, out of late summer and into late autumn. That is the largest single confound in the water quality record and it is almost entirely correctable in analysis — every model in the report carries two day-of-year harmonics — but it is much better fixed at the source. What it does when it is not corrected: water temperature appears to be falling by 1.8 °C per decade, and the same data with the seasonal terms in say it is rising by 0.6 (95% CI 0.31 to 0.85, 1,290 samples). Not a smaller trend, the opposite sign. This is the one item on the list that costs nothing to answer and needs a decision rather than a document: pick a four-week window and hold it.
Blocks: Nothing retrospectively, but every future temperature-sensitive comparison, and the width of every seasonally adjusted interval in the report. Value: high. Costs you: minutes. Refer to it as
dq:sampling-calendar-drift.
37. Would you apply for the full-resolution SEED Greater Sydney tree canopy layers (2016, 2019 and 2022)? It needs an application form and a signed NSW Digital Data Deed Poll, and Council is eligible as a Greater Sydney council.
It is the only remaining route to a network-wide canopy trend. SEED carries the near-infrared band that separates eucalypt canopy from everything else; the Nearmap RGB imagery we have does not, and excess-green calls 1.4% of a closed-canopy tile vegetation. Three epochs is the one thing that could still corroborate or kill the recorded shading trend.
Blocks: A network-wide canopy trend; corroborating or killing the recorded shading rise. Value: high. Costs you: an afternoon. Refer to it as
dq:seed-canopy-deed-poll.
38. Who did the field work, and on which dates, back to 2006? Field diaries, run sheets, timesheets or payroll would all work.
Before 2018 the officer and the period are completely confounded, so we can rule the observer out but never model them. Field officer takes 22-24% of the variance in algal and macrophyte cover — as much as the site itself. A roster turns that from a caveat into a term in the model.
Blocks: Modelling the observer rather than merely ruling them out. Value: high. Costs you: a real search. Refer to it as
dq:field-officer-roster-pre2018.
39. The water quality sheet has an average flow velocity field filled in on about 558 visits. Was that ever measured with a meter, or is it an eye estimate — and if there was a meter, which one and over what years?
We ended up using the free-text water_level box instead, because it has nearly twice the coverage. But a velocity in metres per second, if it is a real measurement, is a better variable than a six-level ordinal judgement, and 558 visits is not nothing. If it was measured, it is worth going back for; if it was estimated by eye, we should say so and stop treating the units as meaningful.
Blocks: Whether flow enters this analysis as an ordinal judgement or as a measured velocity. Value: moderate. Costs you: minutes. Refer to it as
dq:ave-flow-provenance.
40. Does WaterNSW, Sydney Water, NPWS or a university hold a continuous flow gauge inside or near the monitoring network — and if so, can you tell us who to ask for the record?
Flow state recovered from the field sheets carries the level at the moment of sampling and structurally cannot carry antecedent discharge, time since the last flushing flow, or rising-versus-falling limb. Those three are what a gauge adds. Samples taken after rain score lower — 0.042 lower per e-fold of 30-day rainfall — and the recovered flow state does not absorb that, which is the clearest sign something hydrological is missing.
Blocks: Antecedent discharge, flushing history and limb — the three flow variables the field sheet cannot carry. Value: moderate. Costs you: minutes. Refer to it as
dq:stream-gauges.
41. Do the pre-2012 field sheets carry a time of day? The water quality database records a sampling time from 2012 onwards but not before.
Diurnal water temperature variation in a shaded headwater stream is 2-5 °C, and the one water quality trend that survives everything else is about half a degree Celsius per decade averaged over the year. If the hour of sampling drifted over the record it would confound that trend, and at the moment nobody can check the first fourteen years. Only worth the search if the trend gets contested.
Blocks: Testing whether the hour of sampling drifted. Value: moderate. Costs you: a real search. Refer to it as
dq:sample-time-recover-pre2017.
2.5.5 Stormwater
42. Where did the Leura Falls monitoring data end up? Sites E151 and E195-E199, six sites monthly from June 2014 to July 2017, plus rain-event inlet and outlet sampling of five treatment systems with measured pollutant reductions per system.
Not one E-code site exists in either database. This is the only dataset anywhere that could say whether an individual treatment device works — which is a different question from the network-level one the commissioning dates would unlock, and arguably the more useful of the two. It is your own data, recently collected, and somebody knows where it is.
Blocks: Evaluating any treatment system individually; the only per-device performance evidence in existence. Value: transformative. Costs you: an afternoon. Refer to it as
dq:leura-falls-monitoring.
43. When was each treatment asset built? The specific ask: a OneCouncil / CiAnywhere extract filtered to SWDR-GPTS and pit types 07, 08, 11 and 12, returning the FA450 financial-asset commissioning date — plus the DA and construction-certificate records, which have never been searched, and whoever populated “EConst. Date” on 28 SQIDs.
Your central question is whether the stormwater works have improved waterway health, and it cannot be answered because nothing reliably records when the works happened. This would take the evaluation from three usable comparison units to something near sixty, and the smallest change it could detect from more than half a score point to about a fifth of one. Be clear about what it does not do, though: dates make the question askable, not answerable. The monitoring design decides the rest, which is the next item but one.
Blocks: The evaluation question Council asked for; the unexplained 2014.4 changepoint in the health trend. Value: transformative. Costs you: a real search. Refer to it as
dq:sqid-commissioning-dates.
44. Which catchment — or at minimum which upstream drainage node — does each treatment asset serve? It exists nowhere except for the seven Leura Falls systems.
Without it, “downstream of a treatment device” cannot be defined at all, so no monitoring site can be paired to any asset and no asset can be evaluated regardless of how good its date is. This is the quieter half of the commissioning-date problem and it is just as binding.
Blocks: Pairing any monitoring site to any treatment asset. Value: transformative. Costs you: a real search. Refer to it as
dq:sqid-served-catchment.
45. What do these fields actually record — build, handover, capitalisation, load date or revaluation batch? Construction Date in Civica, AssetPit.Year, tblAIMOpenChannel.Datecreate and Effectdate, and Asset_Audit.OC_CommDate.
OC_CommDate is the largest populated date field we have — 204 of 254 treatment assets — and whether it is usable or junk turns entirely on this. Our reading is that it is a year-end capitalisation stamp: it has only eleven distinct values, every one falls on 29 or 30 June, 169 of them are 2016-06-29, and it disagrees with all five independently documented construction dates. But that is inference, and somebody in Assets can just tell us.
Blocks: Whether 204 stored dates are usable. Value: high. Costs you: minutes. Refer to it as
dq:civica-date-semantics.
46. Given that a four-site comparison cannot detect anything, which do you want — evaluate the works on a response that varies less between sites, pool several works programs to reach the site count, or accept in advance that the evidence will be descriptive and say so?
Detecting half a score point needs roughly twelve treated and twelve control sites, sampled twice a year, for six years either side. At four per arm it is unreachable at any amount of sampling — the floor is 0.73 score points. The Leura Falls difference-in-differences returned +0.11 with a confidence interval from -0.36 to +0.58: a null that carries no information at all. Deciding this in advance is what stops a four-site study being run and then reported as “no effect”.
Blocks: Whether the next evaluation is designed to answer something or designed to return a null. Value: high. Costs you: an afternoon. Refer to it as
dq:stormwater-eval-design-decision.
47. Is there a list of the treatment assets that are not pits — rock-lined biofilters, raingardens, bioretention basins, sediment basins, trash racks — and GIS layers for them? They are largely invisible to the asset system.
We have a partial list already: sqid_dates.csv carries 315 treatment assets, including biofilters and rock-lined channels picked up from your 2021 Stormwater System Maintenance Program spreadsheet that are not in the GIS pit table. What is missing is a register that says the list is complete, so we cannot tell whether the gap is five assets or a hundred.
Blocks: Completeness of the intervention set. Value: moderate. Costs you: an afternoon. Refer to it as
dq:non-pit-assets.
48. Where would a decommissioning or major-maintenance date be recorded — is there an asset-disposal workflow in Civica/OneCouncil at all? The Disposed extract returns zero SQIDs, which we read as “none are captured” rather than “none have ever been removed”.
Same mechanism as the commissioning dates and smaller: without them we cannot tell whether an asset was actually operating during the monitoring period it is being credited with.
Blocks: Whether an asset was operating during the period it is credited with. Value: moderate. Costs you: an afternoon. Refer to it as
dq:sqid-decommissioning.
49. Does a fine-grained stormwater sub-catchment layer exist, or an imperviousness layer to go with it? Data\Environment on the corporate GIS share is ACL-denied to us and its orphaned label layers point straight at sub-catchment polygons — could we be given read access to it?
It is what would let the drainage area served by an asset be defined at anything finer than the delineated site catchment, which is currently the only unit available and is far too coarse for a device serving a few streets.
Blocks: Defining served catchments at a useful scale. Value: moderate. Costs you: an afternoon. Refer to it as
dq:stormwater-subcatchment-layer.
2.5.6 Wetlands
50. Is there a second wetland anywhere in the LGA — or a neighbouring one — that could serve as a reference site? One more would change what the wetland rating is able to say; two would change it completely.
Ingar Dam is the only ReferenceWetland in the network. A percentile band drawn from a single site cannot support a condition claim, so the wetland rating is a ranking among the nine sites rather than an assessment against wetlands known to be in good shape. This is the structural reason a wetland rating cannot mean what a stream rating means, and it is the only item on this list that would fix it.
Blocks: Any condition claim for a wetland; the two-anchor band construction the streams use; wetland-specific desirable water quality ranges. Value: transformative. Costs you: an afternoon. Refer to it as
dq:wetland-reference-set.
51. Would you be able to run two or three edge samples at candidate reference wetlands over one season? That is the smallest piece of new fieldwork that would change what the wetland rating can say.
A regional wetland reference set does not need to be large to be transformative here, because the current set is one site. Three wetlands sampled twice each would give a benchmark distribution to score against and would let the wetlands be rated the way the streams are, rather than only ranked against each other. If NPWS or a neighbouring council already sample wetlands to a comparable protocol, their data would do the same job for nothing.
Blocks: Every condition claim about a Blue Mountains wetland. Value: transformative. Costs you: a real search. Refer to it as
dq:wetland-reference-survey.
52. What was behind scoring more than 16 families as a 3 rather than a 5 in wetlands? We have tested it against the Blue Mountains data three ways and cannot find the enrichment it assumes — but if it came from wetlands elsewhere, we have not tested the same claim.
The band is deliberate and documented: a very high family count in a Blue Mountains wetland may indicate excess nutrients. It demotes 33 of the 208 scored edge wetland samples by half a point and moves 6 across the Good/Excellent boundary. Comparing the samples above 16 families with those scoring 5 finds no difference on any of eight water quality indicators; across the whole richness gradient phosphate falls rather than rises; and the tolerant share of individuals does not increase with richness. Our suggestion is to report it as a flag beside the rating instead of folding it into the score — but knowing what the threshold was set from would settle whether that is the right call. The phosphate evidence above is a ratio and a gradient, so it is unaffected by the one thing nobody can pin about that column: whether the readings are phosphate as PO4 or as P, which nothing records and which is worth a factor of three (dq:phosphate-units).
Blocks: The wetland half of any revised rating; chapter 16’s monotonicity test. Value: high. Costs you: minutes. Refer to it as
dq:wetland-nonmonotonic-band.
53. Which samples went into the wetland band table? We have reconstructed it as the 2012–2015 wetland edge samples, but three quarters of those are the six Glenbrook Lagoon points, and we would like to know whether that is what was intended.
The wetland score is a single combined regional comparison rather than the streams’ urban-plus-reference pair, and its percentiles appear to come from 81 edge samples taken between 2012 and 2015 at nine sites. Sixty-one of those 81 are six monitoring points on Glenbrook Lagoon and ten are the reference wetland. If that is the intended construction it is worth saying so in the methods document, because it determines what an “Excellent” wetland is being compared with.
Blocks: Any statement about what a wetland rating means; the wetland half of the revised rating in chapter 15. Value: high. Costs you: an afternoon. Refer to it as
dq:wetland-band-provenance.
54. Site 922, the Wentworth Falls Lake jetty, is classified Urban rather than UrbanWetland, so everything grouped by water body type counts a lake as a stream. Can we reclassify it?
It is water-quality-only, so no health rating is affected, but 510 trigger comparisons from 85 water quality samples currently sit in the stream column. Reclassifying it moves the stream turbidity exceedance figure by about 1.7 percentage points on its own. The macroinvertebrate site for the same lake, 24BWF, is classified correctly.
Blocks: Any water quality analysis grouped by water body type. Value: moderate. Costs you: minutes. Refer to it as
dq:site-922-classification.
55. The desirable water quality ranges are for streams — are you happy for us to report wetland readings without a benchmark until wetland ranges exist, or would you rather we left them out of the report card entirely?
2,500 of the 15,633 comparisons in the trigger table are wetland samples scored against ranges derived from an Appendix 2 site list of 32 stream sites and two wetlands — a stream distribution with a token of wetland in it, not a wetland benchmark. Keeping them moves the headline “% above the desirable range” by three to seven percentage points for EC, salinity, nitrate-N and turbidity, and the wetland subgroup differs far more than that: 59% of wetland alkalinity readings sit above the stream range against 38% of stream readings, and 34% of wetland phosphate readings against 24%. The analysis filters them out and says so; what to do in the published report card is your call.
Blocks: The water quality report card’s treatment of the four wetland water bodies. Value: moderate. Costs you: minutes. Refer to it as
dq:wetland-stream-triggers.
56. Would you like us to derive wetland-specific desirable water quality ranges, the same way the wetland macroinvertebrate bands were derived? It is a day’s work, and it inherits the same one-reference-wetland problem.
There are 644 wetland water quality samples spanning 1998 to 2025, 368 of them carrying at least one of the parameters the desirable ranges cover, which is enough to compute percentiles. What it cannot do is anchor them to wetlands known to be in good condition, so the result would be a wetland ranking on the water quality side to match the one on the macroinvertebrate side. That may still be more useful than comparing lagoon water with a stream range, and it is your call whether it is worth having.
Blocks: Reporting wetland water quality against any benchmark at all. Value: moderate. Costs you: an afternoon. Refer to it as
dq:wetland-desirable-ranges.
57. The wetland family-count column has band 3 ending at 10.99 and band 4 starting at 10.00, so 10.00–10.99 is in two bands. We have read band 4 as starting at 11.00 — can you confirm that is what was meant?
The published column reads “9.0–10.99 or >16” for a score of 3 and “10.00–12.99” for a score of 4. The only reading that makes the column a partition is band 4 beginning at 11.00, and that is what the analysis implements. The methods document’s Wentworth Falls Lake worked example has 12 families, so it does not discriminate between the two readings. It is a typographical defect rather than a conceptual one, but the next person to implement the table could resolve it the other way.
Blocks: Nothing in this report — the choice is made and flagged. It matters for anyone re-implementing the published table. Value: low. Costs you: minutes. Refer to it as
dq:wetland-band-overlap.
2.5.7 Catchments and imperviousness
58. Can we get the drainage network that is missing from the asset register — private property drainage, the Transport for NSW (formerly RMS) highway system and Sydney Water’s assets — or at least be pointed at whoever holds each of them?
Your register reaches a recorded discharge point for 20.3% of its 259 km. The rest ends at an asset boundary, so we have to put the water back on the surface there, and for most of those ends that means a street or a back yard rather than a creek. This is the single largest constraint on anything we can say about connection: we built a traced connected imperviousness and it could only be traced through a fifth of the network, which is why we can say connection did not beat total imperviousness as you can currently measure it, and cannot say anything stronger. It is also what stands between the modelled DCI you asked us for and a measured one.
Blocks: A measured DCI; the connected-imperviousness test; every statement about where urban runoff actually goes. Value: transformative. Costs you: a real search. Refer to it as
dq:complete-drainage-network.
59. How was the “Distance from source (m)” field measured — on what map or imagery, in what year, by whom, and along the channel or as a straight line?
It was probably meant as a descriptive field. It now steers the pour point, and therefore the whole catchment, at 90 of the 120 catchments — 94 have the field recorded, and at four of those no channel was found within reach of the length it gives, so the plain nearest-channel rule was used instead. For most of the 90 the constraint changes nothing: the median catchment moves by 0.08% against a rule that never sees the field. For 15 it moves the area outside the 0.8–1.25 band, and for 11 by more than a factor of two. The extreme is 73BKT, at 7,906 ha with the constraint against 9.5 ha without it.
Blocks: Catchment area at 15 of 120 catchments, and the denominator of every imperviousness figure at those sites. Value: high. Costs you: minutes. Refer to it as
dq:distance-from-source-provenance.
60. Which first_order_catchments layer should we treat as authoritative — the .shp with 1,326 polygons and non-unique IDs, or the .TAB with 526 that carries OBJECTID?
We are currently joining on name because the IDs do not support anything better, and a name join across two layers that disagree on polygon count is wrong somewhere. Every catchment-scale covariate goes through this join.
Blocks: Every catchment-scale join through the first-order catchment layer. Value: high. Costs you: minutes. Refer to it as
dq:first-order-catchments-authoritative.
61. Could we get read access to three folders on the GIS share that currently refuse us — Data\Environment, Data\Images\Model and Data\Assets\Building?
Judging by the names, those are the likely homes of the sub-catchment polygons, the DEM rasters and your own building capture — which is to say several other items on this list. One permission change may close a handful of asks at once.
Blocks: Several other asks on this list. Value: high. Costs you: minutes. Refer to it as
dq:gis-folder-access.
62. Do you need a single imperviousness figure that goes into a planning instrument, or a ranking that decides where work goes first? The two want different things from us, and only one of them is supportable.
Health falls at a steady rate of about 0.03 points of score per percentage point of catchment sealed, right across the range, and a breakpoint model finds no threshold to point at. The ranking of catchments is solid — it is identical whichever imperviousness measure is used — so “protect the least sealed catchments first” is well supported. A numeric trigger is not: on total imperviousness there is no breakpoint in the data, the one figure that can be quoted comes off the modelled connected scale and is a property of that scale’s spacing rather than of the streams (dq:dci-trigger-decision sets that out), and the scale a line would be drawn on moves by a factor of two to three depending on how imperviousness is estimated. If a number is unavoidable for planning reasons, tell us and we will say what can honestly be written beside it.
Blocks: What we recommend in place of the withdrawn 5% DCI trigger. Value: high. Costs you: minutes. Refer to it as
dq:imperviousness-planning-trigger.
63. How far is a typical Blue Mountains roof from the pit it drains to? We have assumed 30 m, with bounds of 15-60 m, and it is an assumption, not a measurement — your drainage engineers will know better than we do.
Median measured DCI moves from 1.5% to 4.7% across that range, and the number of catchments under the 5% boundary moves from 107 to 61 out of 120. The boundary is more sensitive to this one unmeasured distance than it is to the entire modelled-versus-measured imperviousness question we spent a chapter on.
Blocks: The 5% planning threshold, and every DCI figure derived from it. Value: high. Costs you: minutes. Refer to it as
dq:roof-connection-distance.
64. Which point on Glenbrook Lagoon, and which on Wentworth Falls Lake, is the outlet? We have the waterbody polygons and the six lagoon points now share one catchment, but the outlet is inferred as the lowest cell on the polygon rather than known.
Six monitoring points sit on one lagoon and now share a single catchment of 40.5 ha at 25.8% impervious, which is what lets the lagoon into the management priority tables at all. The outlet the catchment drains through is a guess, and moving it moves the catchment. This is a judgement about a lagoon, not something to infer from a flow algorithm, and it is one email.
Blocks: The imperviousness figure behind three lagoon sites in the management priority table. Value: high. Costs you: minutes. Refer to it as
dq:wetland-outlet-decision.
65. Does anything hold invert levels, or upstream and downstream node ids, for the stormwater pipes? AssetPipe has neither, and Depth is one value per pipe so it cannot give a gradient either.
Without them the network cannot be directed from its own attributes, so we inferred the direction: pipe ends within 1 m are the same junction, and every junction drains by the shortest path to its component’s outlet, with the 12 m surface breaking ties. Those are choices, and either of these fields would remove them outright. If they live in a maintenance or design system rather than in GIS, that is just as good.
Blocks: The two modelling choices behind the drainage burn. Value: high. Costs you: an afternoon. Refer to it as
dq:assetpipe-invert-levels.
66. You asked for %DCI so it could go into a planning instrument, and we cannot give you a defensible numeric threshold. What do you want to do instead — and how should a band boundary apply across catchments that differ in size by a factor of thirteen thousand?
The analysis supports ranking catchments and does not support a number. On total imperviousness — the scale the ranking is built on — there is no breakpoint to find at all. The 13.05% figure you may have seen (confidence interval 6.9% to 19.2%) is on the modelled connected scale, which ranks the catchments identically and merely spaces them differently, so a threshold read off it can be a property of the transform rather than of the streams. It is too wide to legislate on even taken at face value. There is also your own banding — under 5% imp, 5-10% im, over 10% imp — and whatever replaces the trigger needs to reconcile with it or replace it explicitly.
Blocks: A planning instrument that is waiting on this. Value: high. Costs you: an afternoon. Refer to it as
dq:dci-trigger-decision.
67. Is roads/AssetRoadSurface.TAB complete enough to use as the sealed road area outright — and does it cover state roads, or only the ones you maintain?
It holds 3,564 surfaced segments with a recorded width and area, about 457 ha of measured seal at a median width of 6.0 m. That is a measurement where we currently have an assumption applied to the whole road reserve, and swapping it in is the single change most likely to improve imperviousness. We have not made it, because it moves every imperviousness figure in the report and the gap where state roads should be would then become a hole rather than a rounding error.
Blocks: The level of every imperviousness figure in the report. Value: high. Costs you: an afternoon. Refer to it as
dq:measured-road-surface-areas.
68. Is there a register of pollution incidents and EPA notices in the monitored catchments? Your own report names a raw sewage leak running into Leura Falls Creek for at least 22 months from April 2016, and a landscaping-supplies discharge that drew a Prevention Notice in September 2017.
Both of those land inside the after window of the Leura Falls before/after case, the only one in the archive, and a control site outside the catchment cannot difference them away. If there are others we do not know about, there are published trends in the health-trend and water quality chapters with undocumented confounds running through them.
Blocks: The Leura Falls before/after case; any trend passing through an undocumented incident. Value: high. Costs you: an afternoon. Refer to it as
dq:pollution-incident-register.
69. Your own per-catchment urban connection factor runs from 0 to 0.99 with a median of 0.46. Does anyone remember what it was based on?
It is recorded as a judgement with no reasoning attached. If it was field knowledge we would like to calibrate our modelled DCI against it rather than merely rank against your labels. If nobody remembers, that is a real answer too and we will say so.
Blocks: Calibrating modelled DCI against your own figures. Value: moderate. Costs you: minutes. Refer to it as
dq:connection-factor-basis.
70. Who produced the published figure of 35 urban sites under 5% DCI, and what did they run? We cannot reproduce it.
The companion figure of 13 reproduces exactly. The 35 matches no combination of connection curve and total-imperviousness bound we can construct; the nearest cell is 29. It is the same class of problem as the rating bands — a published number with no surviving working — and it is better solved by finding the file than by us guessing at it.
Blocks: A published figure that cannot currently be reproduced. Value: moderate. Costs you: minutes. Refer to it as
dq:dci-35-sites-derivation.
71. Your “Typical breakdown” sheet treats driveway and hardstand area as a fixed per-zone multiple of roof area. Was that measured anywhere, or is it a working assumption?
Either is fine; we just need to know which, because it feeds the total imperviousness estimate behind every DCI figure in the report, and as far as we can tell nobody has ever checked it.
Blocks: The total imperviousness estimate underneath every DCI figure. Value: moderate. Costs you: minutes. Refer to it as
dq:driveway-hardstand-ratio.
72. Your own method applies 0.5 impervious to the whole formed road reserve, including unsealed Crown roads and unformed tracks. Was that intended, or is it a side effect of how the layer was built?
Cheap to answer and it matters, because reconciling our imperviousness with your own 2017-18 DCI figures is the one calibration target we have — and because we are making the same assumption at 0.45 for want of anything better. See also the ask about AssetRoadSurface, which would replace both.
Blocks: Reconciling our imperviousness with your 2017-18 figures. Value: moderate. Costs you: minutes. Refer to it as
dq:road-reserve-imperviousness.
73. Could someone re-export urban zones 20170618.csv with the computed fields actually calculated? Connectedness, Connection_factor and DCI_ha all read 0 in every one of the 1,849 rows, which cannot be right — Area_impervious_ha is non-zero throughout.
Probably a save-without-recalculating. It is the only reason we cannot use your own urban DCI component, and it is one re-export.
Blocks: Using Council’s own urban DCI component at all. Value: moderate. Costs you: minutes. Refer to it as
dq:urban-zones-reexport.
2.5.8 The rating system
74. If the spreadsheet cannot be found — were the boundaries deliberately softened or rounded, and does anyone remember the reasoning?
This is the fallback, and it matters as much as the file. If the original author moved a boundary on purpose for a good reason, a mechanical reset would throw that reason away without anyone noticing it had been thrown away.
Blocks: Whether a deliberate judgement is discarded by accident in the reset. Value: high. Costs you: minutes. Refer to it as
dq:band-softening-judgement.
75. Was SIGNAL-SF chosen over SIGNAL 2 for a reason we should know about — a comparability requirement, an agency expectation, an agreement with someone — or was it simply the version in use when the rating was built?
It matters because this chapter’s strongest result is that SIGNAL-SF is structurally the less able of the two to see the change that has actually happened in these creeks: it has no grade for the worms, the chironomid subfamilies and the mites that dominate a degraded sample, and it grades mosquitoes and water striders as though they were clean-water animals. Swapping the factor is a real option, and knowing why the original choice was made would tell us what it would cost you elsewhere.
Blocks: Whether the rating can adopt SIGNAL 2 as a replacement factor. Value: high. Costs you: minutes. Refer to it as
dq:signal2-as-a-rating-factor.
76. Which spreadsheet produced Table 1 of the June 2025 methods document? We want the site list, the date filter, the habitat filter and the percentile calculation behind the 2012-15 urban band boundaries.
The boundaries between Poor, Fair and Good cannot be reproduced from either database under any of six reasonable constructions we tried. You are about to be advised to reset those boundaries, which would move 27 of 77 sites down a class, and that should not happen without knowing why the current ones do not reproduce. It is probably one file on one person’s drive.
Blocks: The band reset recommendation, which is live and costed and should not proceed until this is understood. Value: high. Costs you: minutes. Refer to it as
dq:table1-band-spreadsheet.
77. If the rarefied factors are adopted, will you publish the count of samples that get no rating, by site, alongside the ratings that are published?
Rarefaction cannot score a sample holding fewer animals than the rarefaction depth, so the revised system withholds a rating from 30 samples in the analysis set — and they are not a random 30, because 9 of the 31 Very Poor samples in that same set are among them. That 31 is the Very Poor count over the rated stream edge samples, which is the narrowest of the several populations such a count can be taken over; chapter 1 gives the wider ones and they are not interchangeable. Publishing the surviving ratings without the suppressed count would make the network look better for a reason that has nothing to do with the creeks. We think this is a condition of adopting R2 rather than an optional extra (it is R13), but the commitment has to be yours, because it is your report card.
Blocks: Adopting R2 honestly. Value: high. Costs you: minutes. Refer to it as
dq:withheld-ratings-published.
78. Do you want to adopt the revised rating? It is much freer of laboratory-effort artefact than the current one, but it has not been shown to discriminate better and it cannot rate 30 samples, including 9 of the 31 Very Poor ones in the analysis set.
The honest version of the case. Effort sensitivity — the within-site slope of the standardised score on log sample abundance, comparing samples taken at the same site in the same year — falls from 0.41 standard deviations per natural-log unit to 0.03, an interval that covers zero. That is real, it is the whole point, and the artefact it removes is getting worse as the laboratory counts more animals. Against that: discrimination goes 0.76 to 0.85 on eleven reference sites, the two available tests disagree about whether that is distinguishable from zero, and a temporal holdout reverses it (0.78 for the current system against 0.72 for the revision). Two of the six criteria go slightly the other way — independence and boundary sensitivity. So this is a decision to buy freedom from a known artefact at the price of an unproven improvement and 30 unrateable samples — a trade only you can make, and we are offering it as a suggestion rather than defending it. One scope note on the 31: it is the Very Poor count over the rated stream edge samples, which is the narrowest of the several populations such a count can be taken over; chapter 1 gives the wider ones and they are not interchangeable.
Blocks: The report’s largest recommendation. Value: high. Costs you: an afternoon. Refer to it as
dq:adopt-revised-rating.
79. Could somebody who knows the creeks — and who has not seen the macroinvertebrate data — rank twenty of them by condition for us?
This is the binding constraint on every discrimination result in the report. Your disturbance tiers were built partly from macroinvertebrate health, so testing a macroinvertebrate index against them is circular; the only non-circular yardstick we have is eleven reference sites. Twenty independent rankings would roughly double the independent information available for testing the rating. An hour or two of an experienced person’s time, done blind.
Blocks: Every discrimination result in the rating chapters. Value: high. Costs you: an afternoon. Refer to it as
dq:expert-creek-rankings.
80. Which do you want the network to be good at — the network average, or the individual site rating? You cannot have both on the same budget, and nobody has ever been asked to choose.
Priced in Section 13.8, on the roughly 70 stream samples a year the record currently carries. Spend them on 70 sites visited once each and the network mean has a standard error of 0.100 while one site’s rating has 0.543; spend them on 23 sites visited three times and it is 0.148 and 0.314. One rating class is 1.00 score points wide, so the two products move in opposite directions across most of a class. Our first suggestion — a rolling three-year mean — is free and does not make this trade; buying the same reliability inside a single year does, and nobody has been asked which of the two products the network is for.
Blocks: The monitoring design recommendation. Value: high. Costs you: an afternoon. Refer to it as
dq:monitoring-budget-tradeoff.
81. What, if anything, do you want the public report card to say about phosphate? Our suggestion is to report detectability, label it plainly as a property of the test, and not score it.
Publishing “phosphate met the desirable range in 89% of measurements” — the figure for the 2022-2024 report-card window; over the whole record it is 76% — would report the test kit, not the creek, because most of those measurements are non-detects scored as passes. Detectability is defensible as reporting and is not defensible as scoring. There is reputational risk either way — saying nothing invites the question, and saying the wrong thing is worse — so this one should be a decision you take deliberately.
Blocks: What the public report card publishes for phosphate. Value: high. Costs you: an afternoon. Refer to it as
dq:phosphate-reportcard-decision.
82. Do you want to publish three condition classes, or keep five and print the score with its uncertainty beside the word?
On five classes, a three-sample rating changes the published word 41% of the time, and reaching the protocol’s 20% mark would take about thirteen samples per rating. On three classes the three-sample figure is 23% and the five-sample figure 18%. Neither is comfortable, but five words is the language you have published for a decade and there is a real cost to changing it. Publishing five words off a single sample is not defensible; both of these are. This is suggestion R1b, it is a communication decision about what the public and the councillors are told, and it is yours.
Blocks: What the published rating says. Value: high. Costs you: an afternoon. Refer to it as
dq:rating-class-count.
83. Are you happy for us to keep site-level water quality out of the report card until the sampling calendar and the instrument changes are dealt with? It is the least welcome finding in the chapter and we would rather you decided it than discovered it.
The site-level scan runs 648 site-by-parameter tests and not one survives false discovery rate control, so no creek in the network can be singled out on this evidence. We know officers use site-level water quality to direct field effort, which is exactly why this needs saying out loud rather than by omission. It is fixable — a stable sampling calendar, probe overlap at each changeover, and a defined core set of sites on a fixed schedule would make the question answerable.
Blocks: Any published statement that water quality is changing at a named site. Value: high. Costs you: an afternoon. Refer to it as
dq:site-level-reporting-decision.
84. Chapter 12 found two sites that look as though they are in the wrong disturbance tier, in opposite directions. Our headline discrimination result is measured against those tiers — how much weight do you want it to carry?
This is the analytical half of dq:reference-tier-review, which asks the review question itself. What matters for the rating chapters is the size of the effect. Moving 11BMG to reference and dropping 75BKTR raises the headline AUC from 0.76 to 0.84, so both known errors are working against the rating and the result stands — it is understated, not overstated. But relabelling two sites out of 63 moves the figure by 0.08, and dropping any single reference site moves it between 0.74 and 0.80. One creek’s label is worth several hundredths of AUC, which is the same order as the difference between the two rating systems chapter 15 has to choose between. So AUC against these tiers can say that the rating works; it cannot referee a close contest between two ratings, and we have not asked it to.
Blocks: How much weight the discrimination results can bear. Value: high. Costs you: an afternoon. Refer to it as
dq:tier-errors-calibrate-the-test.
85. When two adjacent percentiles land on the same number — which happens on the integer factors whenever the calibration set is small — would you rather we widened the window until they separate, or collapsed the tied scores into one band and said so in the published table?
Two of the four factors are integer counts and a third is a ratio of two counts, so two adjacent percentiles can be equal. When they are, the band between them is an interval no sample can occupy and the factor can never return that score: a fifth of the scale disappears without anything erroring. One of the six band sets in this chapter has exactly this — the 2015-2019 EPT-family column. The generator now detects it. What it should then do is a policy question rather than a statistical one, and it is yours.
Blocks: Recommendation 9, and Test 6 of the protocol in chapter 16. Value: moderate. Costs you: minutes. Refer to it as
dq:band-tie-policy.
86. Was the 2013-14 fire season taken into account when the 2012-2015 bands were set — and if not, would you want the burn extent of each calibration catchment recorded with the next band table?
The published bands were calibrated on 2012-2015, and the Springwood, Winmalee and Mount Victoria fires burnt from 17 October 2013, inside that window. Of the 43 reference samples the published reference column rests on, 21 come from catchments that burnt over 90% of their area and 11 were taken after the fires. That may be fine — arguably it makes the benchmark more representative of a fire-prone landscape — but nothing in the methods document says it happened, so nobody scoring a creek against those percentiles can tell. It is also why “resetting the bands would bake the 2019-20 fires in” is not quite the right objection: the thing being replaced has a fire in it too.
Blocks: Interpreting any comparison between the published bands and a reset one as a comparison against a clean pre-fire baseline. Value: moderate. Costs you: minutes. Refer to it as
dq:calibration-window-fire-record.
87. Should the report card be built from the macroinvertebrate program’s stream samples, from every stream water quality sample including the recreational program, or from both reported separately?
Both are defensible and they give different answers. On 2022–2024 stream samples the pH pass rate is 67% on the macroinvertebrate program and 49% on all stream samples — a 17-point gap on the same period, the same waterbody type and the same published ranges. The recreational program samples swimming holes rather than the monitoring network, so the two sets genuinely differ. Whichever you choose, the report card has to say which, or it has published an ambiguous number.
Blocks: Every pass rate on the report card, and whether two years of it are comparable. Value: moderate. Costs you: minutes. Refer to it as
dq:report-card-sample-set.
88. Would you rather we reported the edge-only series as the headline trend, with all-habitat as a sensitivity check? All 338 riffle samples are scored against bands derived from edge samples, and they score 0.256 higher than the paired edge sample.
That is 16% of the record and nearly 40% of the first decade, so it is not a footnote. Worth saying that the bias runs the helpful way: removing it makes the improvement larger, not smaller. This is a presentation decision and it is yours, not ours.
Blocks: Which series the report calls the headline trend. Value: moderate. Costs you: minutes. Refer to it as
dq:riffle-band-decision.
89. If the revision is adopted, how would you like a creek that produces no rateable sample to appear in the snapshot — as “very few animals found”, as a blank, or as something else?
Rarefaction cannot rate a sample of eight animals, and the current system gives it a word anyway. Withholding is the honest arithmetic, but it removes bad news selectively — the samples too small to rate are overwhelmingly the ones the current system calls Poor or Very Poor — so a snapshot that simply omits them makes the network look better for a reason that has nothing to do with the creeks. Our suggestion is to publish a count of suppressed ratings by site, and to label the result “very few animals found” rather than “insufficient sample”, because that is what it means. But the wording will be read by the public and it is yours.
Blocks: Recommendation 13, and whether the revision can be published without making the network look artificially healthy. Value: moderate. Costs you: minutes. Refer to it as
dq:withheld-rating-label.
90. Everything we can test is either a single sample or a ten-year site mean. What you publish is a site-year word. Is there any use of that word we should be evaluating directly?
A ten-year site mean averages away most of the noise, so the discrimination figures in chapter 13 are an upper bound on what the published product achieves rather than an estimate of it. We cannot fix this by re-analysis — a site-year comparison on eleven reference sites would be hopeless — but it changes what the phrase “the rating discriminates well” is entitled to mean. If the site-year word drives a specific decision (a works priority list, a report to councillors), tell us which, and we can test the rating against that use rather than against the tiers.
Blocks: How strong a claim the discrimination results support. Value: moderate. Costs you: an afternoon. Refer to it as
dq:rating-grain-mismatch.
2.5.9 Site locations
91. A handful of sites carry a stored lat/lon and an easting/northing that point to different places — 81NFB by 580 m, and 35GFB, 18BKT, 80BMG and 41NWL by over 170 m. Which of each pair is right?
We would rather you adjudicated these than us: a wrong pour point is a wrong catchment, and the catchment is what everything else in the report hangs off. U39 is a separate case — its northing is six digits (626800) where it should be seven, and we have reconstructed it as 6262800. Tell us if that is wrong.
Blocks: Catchment delineation at the affected sites. Value: high. Costs you: minutes. Refer to it as
dq:coordinate-pair-conflicts.
92. Were either of the Hazelbrook bifenthrin locations (87GHZ, 88GHZ) ever sampled before 2023, perhaps under an older site code? Even one visit would do.
There are two bifenthrin contaminations in this archive, eleven years and seventeen kilometres apart, and both are missing the same thing. The 2012 Jamison Creek kill has a control and two impact sites sampled together from five days after the kill was found out to 2024, which is the strongest thing here — but the control and the near impact site were both established after the event, so it has no “before” either (chapter 17). Hazelbrook is the same shape and smaller: the incident is recorded, the sampling only starts afterwards, so the pair is a description rather than a test. It is also the one where a pre-incident sample might plausibly still exist under an older site code, which is why the ask is here and not there. Fifteen minutes of somebody’s memory, or a look at the site register.
Blocks: A real statistical test on the Hazelbrook comparison. Value: high. Costs you: minutes. Refer to it as
dq:hazelbrook-pre2023.
93. Eight sites have two or three plausible positions on the map. Which one is right? Start with 27GLNR — it is a reference site and its two candidates are 1.6 km apart.
The eight are 27GLNR, L4, L13, M23, P7, U13, U39 and U44 (SITE-SUBSETS.md is the register). For seven of the eight the rival positions are between 700 m and 2.1 km apart, which is far enough to give a materially different catchment, so we have withheld them rather than average them — an averaged position is a place nobody sampled. U39 is the eighth and a different problem: its register northing is 626800, six digits where a seven-digit MGA northing belongs. For most of these the register easting/northing picks one point and the creek name picks another, and that is not our call to make.
Blocks: Restoring 73 macroinvertebrate samples and 33 water quality samples to every catchment-scale analysis; two of the seventeen reference-tier sites. Value: high. Costs you: an afternoon. Refer to it as
dq:site-coords-ambiguous.
94. Do the written site descriptions still exist for the ambiguous sites — the “Rian’s Pool, adjacent to the oval” ones? Who would have kept them?
A description like that is the only thing that lets a position be picked off imagery or a map. Without one, the choice between two candidate points cannot be made at all, by us or by anyone else. This is the enabling document for the question above rather than an ask in its own right.
Blocks: Resolving the eight ambiguous positions. Value: high. Costs you: an afternoon. Refer to it as
dq:site-written-descriptions.
95. Are U42 and 07GBH (12 m apart), P2-M6 and 74EBB (24 m), and 32EWD and 32.2EWD (98 m) three sites recorded under two codes each, or six distinct sites?
As pour points each pair is one catchment, so the distinction matters for how the samples are treated: two codes on one site means the samples are not independent, and it also moves the published site counts. One of the three pairs, 32EWD and 32.2EWD, is also on the relocated-site pairs table in chapter 1 (Table 1.8), where the .2 code convention reads it as one site moved to a different reach rather than as one site coded twice. Those are two different answers to this question, and the relocation reading is the register’s convention rather than a confirmed fact, so answering this settles both.
Blocks: Site counts, whether the paired samples are independent, and whether the 32EWD relocation is a relocation at all. Value: moderate. Costs you: minutes. Refer to it as
dq:duplicate-site-codes.
96. Fourteen pairs of site codes in the register are a relocation — 09BBH to 09.2BBH, 59BLA to 59.2BLA and twelve others. Do you want each pair reported as one continuous series, or as two?
A relocated site is a different reach, so we have kept them separate and flagged the relationship rather than splicing them. But you are the ones who will read the trend for Centennial Glen Creek, and if a break in the series at 2024 is going to confuse the reporting, that is worth knowing now. Either answer is fine; what we would rather avoid is the choice being made silently.
Blocks: Nothing — but it should be a decision you took, not a default. Value: moderate. Costs you: minutes. Refer to it as
dq:relocated-site-splice.
97. 77EWD is recorded at 530 m, which puts it in the higher band by the methods document’s own under-500 m rule, but Appendix 2 assigns it to the lower band. Which one did you intend?
It changes that site’s conductivity and salinity desirable ranges and nothing else. We use the altitude-derived band and carry the Appendix 2 value alongside it so the choice can be tested either way.
Blocks: One site’s conductivity and salinity exceedance rates. Value: moderate. Costs you: minutes. Refer to it as
dq:site-77ewd-altitude-zone.
98. Can you confirm the town and creek behind the ambiguous codes? LN is used for both Lawson and Linden, C02 turns up in two coding schemes 2.9 km apart, and 32EWD is Garnett Dam but is typed as a stream in the register.
Anything joined on site code inherits these. The 32EWD one has a second consequence: a waterbody typed as a stream is scored against the stream trigger values and the edge band table, neither of which was derived for it. It also has two names in this report — Garnett Dam here, and Garnett Creek in the relocated-site pairs table (Table 1.8), which takes the name from the waterway column of the register. Whichever is right, one name would be better than two.
Blocks: Any join keyed on site code; the stream / wetland habitat split. Value: moderate. Costs you: minutes. Refer to it as
dq:site-code-decoding.
99. Is there a crosswalk between site codes and the geometry in the five bioindicators/ sampling-point layers and Firesites? None of them carry a site code.
These are layers you already hold and we already have copies of. Without a code on the geometry they cannot be joined to anything, so they are currently unusable — which is a shame, because they are the only spatial record of several sampling programs.
Blocks: Making several spatial layers already in hand usable at all. Value: moderate. Costs you: an afternoon. Refer to it as
dq:site-code-crosswalk.
100. Methods Appendix 2 lists a slightly disturbed site as 38.2NHV; both databases spell it 38.2NVH (Glenbrook Creek tributary, Valley Heights). Which is right?
We have assumed the document has the typo and mapped it to the database spelling, because 38.2NHV appears nowhere else. It matters because the slightly disturbed tier is half the population the desirable ranges were derived from.
Blocks: Nothing — but a one-word confirmation closes it. Value: low. Costs you: minutes. Refer to it as
dq:appendix2-site-code-typo.
101. Are there intact copies of AllMAcroSites22_8_18.TAB (its .DAT sibling is missing) and GISRRatings2018.TAB?
Low priority — we used the .xlsx twins instead and they were fine. It would give a fifth independent lineage to check the recovered coordinates against, which is the only reason it is on the list at all.
Blocks: A fifth cross-check on the recovered coordinates. Value: low. Costs you: minutes. Refer to it as
dq:intact-site-tab-files.
102. If a position genuinely cannot be recovered for a site, do you want to retire it formally — or leave it on the books?
Either answer is fine. What we would like to avoid is a site quietly dropping out of the analyses because nobody ever decided, and then somebody noticing in five years and asking what happened to it. It should be a decision you took, and the report should record it as one.
Blocks: Nothing analytically — but the report should record a decision, not a default. Value: low. Costs you: minutes. Refer to it as
dq:retire-orphan-sites-decision.
103. Are the original field GPS records still anywhere? Stored precision runs from 1 m to 500 m — twelve sites are rounded to 100 m and two to 500 m.
A 500 m rounding cannot choose between candidate points 700 m apart, so several of the ambiguous sites are ambiguous only because precision was thrown away on the way into the database. Going forward the fix is simply to store what the handset gives you.
Blocks: Resolving the ambiguous positions. Value: moderate. Costs you: a real search. Refer to it as
dq:site-gps-precision.
2.5.10 Aerial and satellite imagery
104. Which imagery do you want pulled before the NSW Imagery Hub Planet licence lapses at the end of FY2026-27, and who is going to pull it?
This is the only hard deadline anywhere in the report. Planet imagery reaches you through the NSW Imagery Hub and that licence is funded only to the end of FY2026-27; the Nearmap vertical archive over the LGA (18 surveys, January 2010 to November 2025) sits on your separate Nearmap subscription and is available now, and Nearmap AI Packs, Nearmap 3D and Planet Planetary Variables are not. The obvious candidate is a catchment-scale imperviousness time series, which would turn imperviousness from a single modelled snapshot into a covariate that changes over the record — but the decision is what you want, and the answer only has to arrive before the money stops. Chapter 18 Section 18.6 has what is and is not covered.
Blocks: Any retrospective imagery-derived covariate. After FY2026-27 the same imagery has to be bought. Value: high. Costs you: an afternoon. Refer to it as
dq:imagery-before-licence-lapses.
2.5.11 Incidents and spills
105. Do you still hold the two datasets that sit beside the 2012 Jamison Creek macroinvertebrate sampling — the OEH bifenthrin laboratory results (8 sites, 19 samples, July 2012) and Robert McCormack’s freshwater crayfish surveys (3 sites, 30 samples, July 2012 to April 2014)?
Your conference paper on the incident reports all three datasets together, but only the macroinvertebrate sampling reached the database we have. The bifenthrin results would give us an actual dose against an actual response, which nothing else in twenty-six years of macroinvertebrate monitoring can offer, and the crayfish surveys track the species the kill was noticed by. Both would turn a well-documented incident into a quantitative exposure-response series on one of your own creeks.
Blocks: Any dose-response reading of the 2012 contamination, and any use of the incident to calibrate how long these creeks take to recover. Value: high. Costs you: an afternoon. Refer to it as
dq:jamison-2012-companion-data.
2.5.12 The site register
106. When were the reference and ‘slightly disturbed’ tiers last looked at, and is there a record of who set them and on what evidence?
The tiers are a lookup, not a measurement, and two cases in chapter 12 show them coming apart from the data. 75BKTR (Reedy Creek) is tiered reference and both an external national model and your own scores put it below reference condition. 11BMG (Megalong Creek at Narrow Neck) is tiered urban, and an external model scored the reach beside it band A on 18 of 23 runs while your own 17 samples there average 4.24 — its catchment is 1.2% impervious, inside the range of your reference catchments. Neither is necessarily wrong, because a tier is a judgement about the catchment. But if there is a review cycle we should know its date, and if there is not, that is worth saying in the report.
Blocks: Knowing whether the reference set is the current best judgement or an inherited one. Value: high. Costs you: an afternoon. Refer to it as
dq:reference-tier-review.
107. K40, K43, K53, P7 and W14 are classified Reference in the database but are not in methods Appendix 2 — were they ever reference sites, or is that classification inherited from something else?
They hold 33 macroinvertebrate samples between them, and they are in every count that reads the database classification and in none that reads the methods document. That is why “the reference set” is seventeen sites in some places in this report and twelve in others. Either answer is fine — we just need to know which. Four of the five already have a working position recovered from their stored easting/northing (K43 and K53 confirmed, K40 and W14 probable) and a delineated catchment; only P7 has neither, so P7 is the only one whose location we still need.
Blocks: Reconciling the reference-site counts, and any use of these five sites’ samples in a band derivation. Value: moderate. Costs you: an afternoon. Refer to it as
dq:reference-historic-five.
2.6 Changes to what gets recorded from here on
None of these recovers anything already lost. They are recording changes, and they pay from the next visit onward and not one day earlier. They are on the list because between them they would have prevented most of the questions above it.
1. Would you add a pick-count and subsample-fraction field to the laboratory sheet from now on, and flag any count that was scaled up from a subsample?
The cheapest high-value change available, and it closes the largest confound in this report for everybody who comes after us. It also fixes a second problem: 24% of counts between 100 and 499 are exact multiples of ten, so a count of 500 does not have anything like the precision of a count of 5, and nothing currently records which is which.
Blocks: Nothing retrospectively. Everything from the next sample on. Value: transformative. Costs you: minutes. Refer to it as
dq:pick-count-field.
2. Would you add a per-parameter “below detection limit” tick box to the laboratory sheet, and a < prefix to the value?
The single highest-yield form change on this list. This one box would have prevented 101 mis-stored coliform records and a published trend that turned out to be an artefact, and it would have made the biggest question on the whole master list unnecessary.
Blocks: Nothing retrospectively. Value: high. Costs you: minutes. Refer to it as
dq:below-detection-tickbox.
3. Do you want to complete flow velocity, discharge, average depth, wetted width and cross-sectional area on every visit — or take them off the sheet altogether?
Flow is probably the most important physical driver of a macroinvertebrate community and it appears nowhere in this report, because those fields are populated on roughly one visit in five. A field filled one visit in five is zero of a dataset plus the cost of collecting it. Either is a defensible answer; the current state is the one that is not.
Blocks: Nothing retrospectively. Value: high. Costs you: minutes. Refer to it as
dq:flow-fields-complete-or-drop.
4. Would you replace the free-text officer field with a staff dropdown, on both the field and laboratory sheets — and reconcile the names already in there? There are eight labels for six people.
Near-zero cost. In ten years this is what separates a change in the creeks from a change in who is looking at them, and at the moment there is nothing before 2018 to do that with.
Blocks: Nothing retrospectively. Value: high. Costs you: minutes. Refer to it as
dq:officer-dropdown.
5. Would you record the clock time of sampling on every visit — and make sure it survives into the database, which the current one does not?
Diurnal water temperature variation in a shaded headwater stream is 2-5 °C. The one water quality trend that survives everything else in this report is about half a degree Celsius per decade averaged over the year. A slow drift in the hour of sampling would confound it completely and nobody could check, because the time is collected on the sheet and then thrown away on the way in.
Blocks: Nothing retrospectively. Protects the only surviving water quality finding from here on. Value: high. Costs you: minutes. Refer to it as
dq:sample-time-through.
6. Would you keep the free-text field notes and give them more room on the sheet, plus tick boxes for “result estimated”, “plate crowded” and “too many to count”?
The notes are the best thing in the dataset and they cost nothing. They are the only reason the coliform detection limit is now documented at 10 rather than guessed at 1. More room and three tick boxes would make them better still.
Blocks: Nothing retrospectively. Value: high. Costs you: minutes. Refer to it as
dq:structured-field-notes.
7. Would you complete all eight substrate classes, and the three vegetation strata, on every visit from now on? Both are already on the sheet.
Fine sediment is the strongest predictor of sensitive rare-family richness in this report, and the silt and clay boxes are the ones most often left blank. The three vegetation strata are recorded on about half of visits. A field that is filled some of the time cannot carry a trend, and neither of these costs anything to fix — they are already printed on the form.
Blocks: Nothing retrospectively. Any fine-sediment or riparian-structure trend from here on. Value: high. Costs you: minutes. Refer to it as
dq:substrate-classes-complete.
8. Would Assets backfill Construction Date for the treatment assets — it is populated for 38 of 223 — or confirm that it never will be? And start recording disposals, and enter the treatment assets that are not pits?
The draft WSUD register in TRIM lists about 260 devices with zero dates and names “Construction Date” as an attribute still to be collected, so this is already known internally. The whole evaluation question depends on it being fixed once. “It never will be” is a genuinely useful answer, because it tells us to stop designing around a register that is not coming.
Blocks: Nothing retrospectively. The evaluation question from here on. Value: high. Costs you: an afternoon. Refer to it as
dq:asset-register-completeness.
9. Would you write down a protocol for the habitat observations, and set up a photo point at each site?
Ten of the thirteen habitat variables drift within a fixed site, and there is no written protocol behind any of them. Until there is one, recorded shading has to be read as a state variable — good for comparing reaches, useless for comparing years — and a rise in it must not be read as improvement. A photo point makes the next twenty years auditable for the price of a stake in the ground.
Blocks: Nothing retrospectively. Protects the whole physical habitat block from here on. Value: high. Costs you: an afternoon. Refer to it as
dq:habitat-protocol-photopoints.
10. At the next probe changeover, would you run two field rounds with both instruments before retiring the old one?
Two days of fieldwork protects the next twenty years of data. The absence of exactly this at the last two changeovers is why turbidity and dissolved oxygen have been abandoned in this report.
Blocks: Nothing retrospectively. Prevents a third unsplittable instrument step. Value: high. Costs you: an afternoon. Refer to it as
dq:probe-overlap-next-change.
11. Would it be feasible to record values once, on a tablet at the creek, with each parameter’s plausible range built in — so that a DO% of 7.8 beside a DO of 8.8 mg/L asks a question at the creek rather than twenty years later?
This removes an entire error class rather than a particular error. Every single item in the transcription register is a hand moving between a clipboard and a keyboard, and the post-2020 probe data — which writes its own file — has none of them at all. We do not know what a system change costs you, which is why this is a question rather than a suggestion.
Blocks: Nothing retrospectively. Value: high. Costs you: we do not know. Refer to it as
dq:tablet-form-at-creek.
12. Would you replace the free-text water_level box with a six-option dropdown — no flow, low, low-moderate, moderate, moderate-high, high?
Nothing is lost retrospectively: the parser recovered 1,003 of 1,005 rows, which is a testament to how consistently people wrote in that box. But the next seventeen years should not need a parser, and the six options are the ones your own field staff already use.
Blocks: Nothing retrospectively. Value: moderate. Costs you: minutes. Refer to it as
dq:flow-state-dropdown.
13. Would you store the imperviousness figure for each sub-catchment in the database, alongside the health results? Neither “DCI” nor “impervious” appears anywhere in either database at present.
Your own sub-catchment classification matrix — the thing that decides whether a creek is reference, slightly disturbed or urban — cannot be reproduced from your own records, because the number it turns on is not in them. We rebuilt it from public data and it took weeks. One column would mean nobody ever has to again, and it would let the tier assignment be audited rather than trusted.
Blocks: Nothing retrospectively. Reproducing your own tier assignment from your own records, from here on. Value: moderate. Costs you: an afternoon. Refer to it as
dq:imperviousness-in-database.
14. Would it be possible to flag a sample in the database when something has happened to the creek — a spill, a contamination, a fire, works in the channel — rather than leaving it to a free-text comment or to nothing at all?
The July 2012 bifenthrin contamination of Jamison Creek destroyed the macroinvertebrate community for three months and is completely invisible in the database. Twenty-four samples across the impact and recovery period carry no marker of any kind, and anyone analysing the Jamison series without knowing the history would read a catastrophic contamination as a poor run of monitoring results. The 2023 Hazelbrook incident is recorded, but only as two sentences in a site note. This one recovers nothing already lost - it is about the next incident, not this one.
Blocks: Nothing retrospectively. Going forward, it is the difference between an event that can be analysed and one that has to be remembered. Value: moderate. Costs you: we do not know. Refer to it as
dq:flag-incidents-against-samples.
2.7 Asks that are not addressed to you
These go to DCCEEW and NPWS, Sydney Water or the Greater Sydney Landcare Network rather than to you. They are here so the whole set is in one place, not because anyone expects you to chase them.
1. (DCCEEW, not you) Which unit was TKN actually reported in? The per-row Units field and the t_WaterQualityTypes lookup both say µg/L, but the stored values are only sensible as mg/L. And is -999 in O/E50 and band a null?
A thousandfold error waiting to happen, in a dataset the report leans on as its only external yardstick. Cheap to confirm and expensive to get wrong.
Blocks: Any use of the external reference set as a yardstick. Value: moderate. Costs you: minutes. Refer to it as
dq:ausrivas-units.
2. (DCCEEW / NPWS) NPWS records a burn in 22 catchment-seasons since 2021-22 where FESM maps nothing — almost all of them prescribed. Which product should we believe?
A zero on the severity raster is not always a zero fire, and every fire covariate in the catchment, health-trend and community chapters inherits the answer. Prescribed burns are the obvious explanation — too cool or too patchy for the severity mapping to pick up — but that needs confirming rather than assuming.
Blocks: Every fire covariate in the report. Value: moderate. Costs you: an afternoon. Refer to it as
dq:fesm-vs-npws-fires.
3. (DCCEEW / NPWS) The fire layer stops naming fires before 2002, and the largest fire season in the whole monitoring record — 2001-02 — falls on the wrong side of that boundary. Is there an incident register or a season report that would name them?
This is a gap in the source rather than a defect in the layer, and it is worth saying plainly which way it runs. Of the 894 catchment-seasons the NPWS layer records a burn in, 554 carry no fire name at all — but 508 of those are before 2000, and from 2002 on every single burnt catchment-season is named. The whole of the remainder is one season: 46 of the 50 burnt catchment-seasons in 2001-02 have no name. That season is the largest in the record by area burnt across the monitored catchments, larger than 2019-20 and larger than 2013-14, so the one season a reader is most likely to ask about by name is the one the layer cannot answer for. Nothing in the report is wrong because of it: every fire covariate is built from the burn geometry and not from the name, and chapter 4’s reference-catchment table (Table 4.17) says in terms that a blank is a missing name rather than a missing row. What is lost is the ability to tie a burn in a particular catchment to a fire anyone remembers.
Blocks: Naming the fires behind the 2001-02 burns in any table or map, and any account of a particular catchment’s fire history before 2002. Value: low. Costs you: minutes. Refer to it as
dq:npws-fire-names-before-2002.
4. (Sydney Water, not you) Is there a populated sewer extract? YEARLAID, DISUDATE and PLANNUM are empty for all 37,659 features in the one we have. The September 2005 trunk mains amplification REF, the reticulation network, the amplification staging dates and the Winmalee STP discharge history would all help.
This is no longer the coliform question — that break is identical at unsewered reference catchments, so the sewer is not the cause. But nothing else dates the sewer network, and catchment history is thin without it.
Blocks: A catchment-history covariate. Value: moderate. Costs you: a real search. Refer to it as
dq:sydneywater-sewer-history.
5. Would you ask Sydney Water for records dating the original Vale St wetland cells? They were built by Sydney Water in the 1990s rather than by Council, which is why they are not in your register.
One asset, but a load-bearing one: it is the intervention date for the one catchment with a defensible before-and-after case.
Blocks: The one defensible before/after case in the archive. Value: moderate. Costs you: a real search. Refer to it as
dq:sydneywater-valest.
6. (DCCEEW) Is t_AUSRIVASPhysChem available anywhere? It is empty statewide in the public extract.
Only needed if we want to re-run the AUSRIVAS models rather than accept the published O/E50 scores, which is not currently the plan. Low priority, on the list so it is not rediscovered later.
Blocks: Re-running AUSRIVAS models rather than accepting the published O/E50. Value: low. Costs you: an afternoon. Refer to it as
dq:ausrivas-physchem.
7. (Greater Sydney Landcare Network, not you) Streamwatch’s published 1990-2020 spreadsheet already covers Blue Mountains creeks; does GSLN hold the site-level detail and the post-2020 rounds that would sit either side of the 2005 and 2017-2019 discontinuities?
The 1990-2020 spreadsheet is on NSW SEED under CC BY 4.0 and we should fetch that ourselves rather than ask for it; what is worth asking for is what the spreadsheet does not carry. It is the only external record that might independently show whether those two breaks are in the creeks or in the method – your own records cannot explain either.
Blocks: An external check on two unexplained breaks. Value: low. Costs you: an afternoon. Refer to it as
dq:streamwatch-records.
2.8 Decisions about buying things
Purchasing decisions rather than data questions. One of them is time-limited.
1. Do you want to renew the Planet / NSW Imagery Hub licence? It is funded only to the end of FY2026-27 and has no dedicated staff support now — and if it is going to lapse, we should pull what we want before it does.
There are 1,860 clear PlanetScope scenes over the study area since 2017, with two years of pre-fire baseline and 221 clear scenes through 2020. That is enough to turn fire from a yes/no covariate into a post-fire recovery trajectory, which is exactly what the depressed reference benchmark needs. This is the most time-sensitive item on the whole list.
Blocks: Converting the fire covariate into a recovery trajectory. Value: moderate. Costs you: minutes. Refer to it as
dq:planet-licence-renewal.
2. Do you want to buy the Nearmap AI packs and 3D/DSM coverage? You have no entitlement to either at the moment — packs.json returns 403 and there are zero 3D surveys across the LGA.
They would give ready-made canopy polygons, which would save real work. But the SEED canopy layer is the better route and it is free to a Council, so this is a convenience purchase rather than a necessary one.
Blocks: Nothing that SEED canopy would not also unblock. Value: low. Costs you: minutes. Refer to it as
dq:nearmap-entitlement.