Soil DNA survey

Daniel Tang

A student at Langley High School, who spent a semester working on a regenerative farm and then spent considerably longer finding out what was living in its soil.

A semester on the farm

This started as a semester on a working regenerative farm, where the difference between a healthy field and a tired one is something people talk about constantly and measure almost never. Farmers judge soil by how it smells, how it crumbles, how fast water disappears into it — good instincts, built on an ecosystem nobody on the field can see.

The study that came out of it is not about that farm's methods, though for a while it was written up as if it were. The question that actually got asked came from the landowner, and it was narrower and better: there is a weedy corner of the property, and are the plants growing there — the ones you would normally pull — doing anything for the soil underneath them? That is what the thirteen cores can answer.

That gap is what this project is about. A teaspoon of that soil holds on the order of a billion bacteria and thousands of distinct kinds, almost none of which will grow in a laboratory. Reading their DNA is currently the only practical way to take attendance.

White canopy tents in a parking lot under an overcast sky, tables of
                  seedlings and produce beneath them, a few people walking between the stands.
Selling at the market stand, a Saturday morning in June 2025. Most of the semester looked like this, not like a laboratory. The soil question came later.

What the project turned into

Soil samples from under five different plants in the same pasture were sequenced, and the resulting data — along with some earlier analysis of it — became the material for a full re-analysis: a scientific manuscript, a reproducible codebase, and this site.

A good deal of that work was undoing rather than adding. An inherited set of "fungal guild" results turned out to be an artifact of matching non-fungal DNA against a fungus-only database. A headline ratio circulating in earlier drafts mixed two levels of classification and had to be recomputed. A public dataset picked as the best available comparison turned out, on inspection of its actual reads, to contain three different marker genes and no usable control. None of that is exciting to find, but each one changed what the study can claim.

What I took from it was to check the data against itself, and to write down what I could not say. Both the paper and this site carry a section listing what the study does not establish.

A pale grey and white turkey with its tail and wing feathers fanned
                  out, standing on a sunlit gravel farm track.
My neighbours for the semester. This one liked to supervise.

Whose data this is

The soil belongs to the farm, and so does what was found in it. The cores were taken from working pasture with the owners' permission and at their invitation, and the data from them — the sequences, and every result computed from those sequences, including the figures on this site — is the farm's. This project analysed that data. It does not own it, and it is not the farm's published record of its own ground unless the farm chooses to make it so.

The analysis code and the writing are the author's, and so is responsibility for any error in them. Where this site interprets a result — and especially where it says what the study cannot establish — that is the author's reading, not something the farm has said about its own soil.

Looking out from a covered porch at sunset over rolling pasture: a
                  farmhouse to the left, a small log cabin below, a willow, and a track
                  running over the hill under pink and orange cloud.
Looking out over the pasture on an evening in July 2025. Somebody's home and livelihood before it was anybody's study site.

Getting in touch

Questions and corrections are welcome through the project repository. Corrections especially: several of the results here exist in their current form because someone questioned an earlier version of them.

Requests for the underlying data are a matter for the farm rather than for this project, so they are passed on rather than answered here. The analysis code is a separate question, and the repository is the place for it.