1-4
How Do We Reconstruct Ecological Interactions?
Species composition tells us which organisms occur in a community, but ecosystems are shaped by the connections among them. For pollinators, these connections are formed when insects visit flowers, collect pollen and transport plant material across the landscape.
Our research uses the biological traces carried by pollinators to reconstruct these otherwise difficult-to-observe interactions. Genome skimming can identify and quantify the pollen associated with individual bees, while genomic reference data and machine learning can help connect complex biological mixtures to their ecological and geographic origins.





From community composition to ecological connections
A species list identifies the members of a community, but not the links among them. Pollen carried by an insect records the plants it has visited, turning the pollinator itself into a biological sampler of the surrounding landscape.
Can DNA reveal who interacts with whom in a natural community?
Finding the links hidden within a community
Community sequencing can reveal which plants and insects occur in an ecosystem and how strongly they are represented. Yet two communities with similar species composition may function very differently if their species interact in different ways. Reconstructing those connections requires evidence that links one organism directly to another.
Plant–pollinator interactions provide a clear example. Flower visits can be recorded by direct observation, but observations are labour-intensive and may not reveal which pollen an insect actually collects or transports. Pollen recovered from the body, pollen baskets or internal contents of a pollinator provides a more direct biological record of its use of floral resources.

Pollen reveals hidden plant–pollinator interactions: challenges and opportunities
(AI-generated figure)
The difficulty lies in identifying mixtures containing pollen from several plant species and estimating their relative contributions. Microscopic identification requires specialist expertise, may not separate closely related plants and can overlook rare pollen types. PCR-based metabarcoding increases taxonomic coverage, but unequal amplification among plants can distort pollen proportions.
Quantifying the pollen carried by pollinators


We tested whether PCR-free genome skimming could identify and quantify mixed pollen. Known pollen mixtures were assembled across a wide range of species proportions and sequenced directly, without amplifying a selected barcode.
Every plant species was detected across the experimental mixtures. The proportion of reads assigned to each plant was strongly correlated with its proportion of pollen grains, and more than 97% of the taxa were placed within the correct order of magnitude. Even plants contributing only a small fraction of the mixture were recovered. The results showed that genome skimming could provide both the identity and relative quantity of pollen in complex samples (Lang et al., 2019).
The method also worked with the amount of pollen collected from a single honey bee pollen basket. This is important because interaction networks are built from individual links. If the pollen carried by one insect can be identified and quantified, each specimen can contribute a record connecting a particular pollinator to the plants it used.
The resulting evidence is more informative than a simple record of flower visitation. It can distinguish broad dietary use from concentrated reliance on particular plants and can reveal rare floral resources within a mixed pollen load. Applied across many insects, sites and sampling periods, these individual records can be assembled into a quantitative network of plant–pollinator interactions.
Why pollen reads are not automatically pollen counts

Sequence reads must be calibrated to estimate pollen abundance.
(AI-generated figure)
Removing PCR preserves a clearer relationship between pollen abundance and sequence abundance, but the relationship is not identical across plant species. Pollen grains differ in size and structure, and the number of plastid genomes contained within a grain can vary among taxa. These biological differences influence the amount of plant DNA recovered during sequencing.
The 2019 study therefore used known mixtures to calibrate species-specific differences in plastid genome representation. This calibration substantially improved pollen quantification. The result establishes both the value and the limit of the method: genome skimming retains quantitative information, but accurate conversion from reads to pollen proportions requires suitable plant references and biological calibration.
This distinction is important for ecological interpretation. Uncorrected reads may still show which plants were used and provide an approximate measure of their representation. Stronger claims about pollen numbers, interaction strength or resource preference require calibration and careful sampling.
Biological mixtures also record place

Pollen does more than connect pollinators to plants. Because floral communities vary among landscapes, pollen mixtures can carry information about where an interaction occurred.
We explored this principle using honey produced in several regions of China. Honey contains plant-derived DNA from the floral resources used by bees, along with DNA from the bees themselves and other biological components. Shotgun sequencing recovered these mixed genomic signals without requiring prior knowledge of the complete local flora (Liu et al., 2022).
Direct taxonomic interpretation remained limited because fewer than one-fifth of the plant-derived sequences closely matched entries in the public reference database. The biological mixtures nevertheless differed consistently among locations. A neural network trained on these sequence profiles assigned nearly all tested honey samples to their geographic origins, including samples from neighbouring sites separated by only short distances.
This result introduced a complementary way of interpreting ecological mixtures. When reference libraries are incomplete, it may not be possible to name every plant contributing to a sample. Even so, the combined genomic profile can serve as a reproducible signature of the local biological community. Machine learning can recognize this multivariate pattern without requiring a predetermined list of indicator plants.
The method depends on a well-represented training collection. Locations represented by too few samples provided less complete sequence signatures and weaker assignments. Local reference samples therefore remain essential, even when the analysis bypasses full taxonomic identification.
Separating the origin of the interaction partners
A biological sample may contain signals from several levels of organization. Pollen can identify the plants used by a pollinator, while the pollinator’s own genome can reveal the population from which it originated. Distinguishing these layers can help separate locally formed interactions from those involving managed, transported or dispersing organisms.

Our later work developed TraceNet, a deep-learning method that assigns individual organisms to source populations using genome-wide SNP variation. Tests with the Eastern honey bee, the red imported fire ant and chicken populations achieved high assignment accuracy, including cases in which population structure was complicated by incomplete separation or gene flow. Filtering for informative genomic sites further improved performance (Yang et al., 2024).
The value of this approach lies in connecting a query specimen to an established population reference without reconstructing an evolutionary tree each time. It also identifies the genomic sites that contribute most strongly to an assignment, making the basis of the prediction more interpretable. As with the honey-tracing analysis, its performance depends on comprehensive and balanced reference sampling.
For ecological interaction studies, these methods provide complementary information. Plant DNA records the resources used by an insect. The insect genome identifies the pollinator and can, where sufficient population references exist, help establish its geographic origin. The interaction can therefore be examined not only as a link between two species, but as a connection involving particular populations within a defined landscape.

Extending the framework to Churchill
This progression now informs our ongoing research in Churchill. The earlier studies established that pollen genome skimming can recover plant identities and approximate their relative contributions from the pollen carried by individual bees. The current programme applies this principle to a natural Subarctic pollination community.
We are building the necessary reference framework from both sides of the interaction. Insects collected while visiting flowers are identified through DNA barcodes, while local plant specimens and leaf tissues provide taxonomically verified references for the floral community. Pollen associated with individual insects can then be matched against these plant references, converting each specimen into a recorded link between a pollinator and one or more plants.
Repeated across species, habitats and the flowering season, these records will allow us to move beyond inventories of local plants and insects toward a quantitative pollination network. They can show which interactions are widespread or specialized, how floral resource use changes through time and whether different pollinators rely on overlapping or complementary sets of plants.
The Churchill programme also addresses a limitation exposed by the honey-tracing study: public sequence databases are often incomplete for local floras. Building a verified regional plant reference library is therefore not a preliminary technical exercise. It determines how many pollen sequences can be assigned to named plants and how accurately the ecological links can be interpreted.
Machine learning offers a complementary route when full taxonomic assignment remains impossible. Repeated genomic profiles from insects, pollen and sites may allow samples to be compared through their combined biological signatures, while expanding local references gradually converts those signatures into identifiable species and interactions.
From observations to ecological networks
DNA does not demonstrate every aspect of an interaction. Pollen on an insect records contact with a plant, but it does not by itself prove that effective pollination occurred. Some pollen may have been acquired indirectly, and the amount transported does not necessarily equal the amount deposited on a receptive flower.
The evidence is nevertheless stronger than co-occurrence alone. A plant and pollinator found at the same site are only potential partners. Plant DNA recovered from an individual insect records a direct biological association between them. When supported by field observations, specimen metadata and local reference libraries, these associations provide the edges needed to reconstruct ecological networks.
The transition from composition to interaction changes the question asked of biodiversity data. Instead of only determining which species occupy an ecosystem, we can begin to examine how they share resources, how strongly they depend on one another and how these relationships vary across space and time.


Why ecological interactions matter
Ecosystem function depends not only on the species present, but also on the relationships connecting them. A decline in interaction frequency, loss of a key floral resource or shift in resource partitioning may alter pollination before either partner disappears from the community.
Pollen genome skimming provides a means of identifying and estimating the floral resources associated with individual pollinators. Genomic reference libraries connect these records to named species and populations, while machine learning can recognize geographic and biological signatures when complete taxonomic references are not yet available.
In Churchill, these approaches support the transition from documenting Subarctic biodiversity to reconstructing the pollination networks that organize it. The resulting networks provide the foundation for asking how environmental change reshapes interactions and, ultimately, ecosystem function.
Nest question
Ecological networks describe how species are connected today, but those species and their interactions are products of a much longer history.
How can evolutionary relationships help explain the origin and distribution of biodiversity?
This question leads to the next feature:
How Does the Tree of Life Explain Biodiversity?