Soil DNA survey

The community moves along one gradient

The tidiest story this survey found is that ammonia-oxidising archaea are scarce under clover because the clover leaks nitrogen and leaves them less to do. It is a good story, and it fails a test the data can run against it. The association is real; it is just not the archaea answering to the legumes. The whole community moves along a single dominant direction, and the archaeal share is a readout of position on it.

The tidy story

Across the thirteen cores, the archaeal share of the DNA falls where the nitrogen-fixing rhizobia rise. The correlation is strong — Spearman's ρ = , ranking the cores rather than trusting their raw values — and the two-group split between legume and non-legume patches is itself significant (p = ). Read forward, that is a mechanism: fixed nitrogen reaches the soil, the archaea have less ammonia to oxidise, and their numbers fall. It is the kind of result a paper is built around.

So ask the cheap question first. Would a number that strong turn up even if the archaea were not answering to the legumes at all? Here it would, and the rest of this page is why.

The result is real, but it is not specific

A correlation is only evidence for the story told about it if it is specific to the players in that story. So the fixers' correlation with the archaea is ranked against the same correlation computed for every other reasonably abundant genus — all that clear a small abundance and prevalence floor. If the nitrogen mechanism were doing the work, the fixers should sit at the top of that ranking.

They do not. The nitrogen-fixers rank , and unrelated genera — groups with no known part in nitrogen fixation — track the archaea at least as closely, past a threshold of ρ = . A story that singles out the legumes cannot explain why so many groups that have nothing to do with them follow the archaea just as faithfully.

Two duller explanations are ruled out before drawing that conclusion. The first is compositional closure: in a fixed total of reads, a slice as large and variable as the archaeal share (% to % here) pushes almost everything else the other way by arithmetic alone. Re-run on centred log-ratio abundances, which are free of that constraint, the association survives (ρ = , p = ), with genera still past the line. Closure did not manufacture it, so the association is genuine. What fails is its specificity.

What the archaea are actually tracking

That negative result points at something. If dozens of unrelated genera all follow the archaeal share, the simplest reason is that the whole community varies along one dominant direction and the archaeal share is simply a position along it. That is directly checkable. The leading axis of the ordination — the same projection the home page plots — accounts for % of the variation between cores, and the archaeal share tracks it tightly: ρ = , p = .

So the clover correlation and the fescue correlation and a dozen others are not a dozen separate facts about a dozen plants. They are all the same fact: the community sorts along a single axis, the plants sit at different places on it, and the archaea rise and fall with the axis. The legume signal is real because the legumes really do sit at one end — but the mechanism the number seemed to name is not established, because the same number falls out of the community's overall structure without any nitrogen story at all.

And it is not a simple growth-strategy gradient either

The obvious next question is what that dominant axis is. One standing candidate is a life-history gradient: soils are often described as running from slow, resource-poor communities of oligotrophs to fast, resource-rich communities of copiotrophs, and fast growers tend to carry more copies of the 16S gene per genome. If the axis were that gradient, community-weighted copy number should climb along it.

Each core's community-weighted mean copy number was computed from , a published per-taxon database, covering % of reads. It ranges from to copies per genome across the thirteen cores, so there was real variation for the check to find.

It found very little. Copy number does not differ between legume and non-legume cores (p = ), shows no clean gradient along the dominant axis, and barely tracks the archaeal share (ρ = ). The same test rules out the dull artifact it was first run to catch — the archaea-to-bacteria ratio is not an effect of copy number — but it does not name the axis. Whatever organises this community, a plain copiotroph-to-oligotroph reading is not it.

Why this is worth publishing

The mistake this result guards against is a common one. Small studies that compare microbial communities across a handful of conditions — five plants, three treatments, two soils — routinely read a guild's correlation with one condition as evidence for a mechanism tied to that condition. This dataset shows how that can go wrong even when every individual number is sound: the correlation is strong, it is not an artifact of closure, and it is still not specific, because the guild is riding a single community axis that many things move along together.

So the finding is narrower than the tidy version, and it holds up better. There is a real, dominant axis of community structure in this pasture; the plants sit at different points on it; and the striking archaeal numbers are a readout of that axis rather than a demonstrated response to legume nitrogen. Naming what the axis is — which needs the measurements set out under ground truth — is the open question this leaves behind.