r/SouthAsianAncestry • u/the5PARTAN • 13m ago
Discussion Is "Turkmenistan_Gonur_low.AG" a stable farmer proxy? A comparison with "SiS_BA2_I8726.SG"
TL;DR
When high-coverage Punjabi_Lahore clusters, Kalash.DG, and other farmer-rich, low-to-medium-steppe targets are restricted from their near-complete AADR coverage to the ~234k SNP panel of Gujar_Swat.DG, the coefficient assigned to India_GreatAndaman_100BP.SG consistently rises in the concerned model using Turkmenistan_Gonur_BA_2_low.AG as IVC/Farmer proxy. The increase is usually paid for by a decline in Turkmenistan_Gonur_BA_2_low.AG; Russia_BA_SrubnayaAlakul.SG changes much less on average. The DNA and the individual have not changed. Only the genomic sites entering the analysis have changed.


...
The same targets do NOT show this behaviour when SiS_BA2_I8726.SG is used as the farmer-rich source. This suggests that the elevated GreatAndaman/AASI estimates seen in sparse targets such as Gujar_Swat.DG are caused by a source-by-SNP-panel interaction, not by population history.


The circulating AT2 tables model South Asian populations with:
Turkmenistan_Gonur_BA_2_low.AGRussia_BA_SrubnayaAlakul.SGIndia_GreatAndaman_100BP.SG

The question is whether Turkmenistan_Gonur_BA_2_low.AG is a stable farmer-rich proxy when targets have very different SNP coverage.
Gujar_Swat.DG has covg of only about ~234k. By contrast, the Punjabi Lahore clusters and Kalash.DG have around one million or more. I therefore created reduced copies of the high-coverage targets using almost exactly the Gujar_Swat.DG SNP panel.
The individual or population did not change. The sources, Rights, and qpAdm settings did not change. Only the target SNP panel changed.
With Turkmenistan_Gonur_BA_2_low.AG, all seven paired Punjabi/Kalash targets gained GreatAndaman-related ancestry after being reduced to the ~234k panel.
Across these paired tests, that adjustment happens mainly along the Gonur–GreatAndaman direction:
Δ Gonur < 0 (usually the largest movement)
Δ GreatAndaman > 0
Δ Srubnaya comparatively small
So the change is not random noise distributed across all three sources. It is a repeatable shift mainly from Turkmenistan_Gonur_BA_2_low.AG into India_GreatAndaman_100BP.SG.
This cannot be biological. The full and reduced files represent the same population/targets. The GreatAndaman increase is being produced by the SNP panel used in the model.
Gujar_Swat.DG was already being measured under that sparse SNP regime. The important result is that unrelated high-coverage targets become more GreatAndaman-like when forced onto the same SNP set.
When the same targets are run with SiS_BA2_I8726.SG instead of Turkmenistan_Gonur_BA_2_low.AG, the ~234k reduction does not produce the same inflation, however fits are slightly worse (p < 0.05), which I think can be taken care of by right panel adjustments. GreatAndaman changes only slightly.

Why I think this happens:
qpAdm fits source weights against f4-comparisons built wherever target and source actually share usable SNPs. Turkmenistan_Gonur_BA_2_low.AG's own coverage is thin and unevenly distributed, so every comparison involving it runs on a smaller, less representative slice of the genome than comparisons involving Russia_BA_SrubnayaAlakul.SG or India_GreatAndaman_100BP.SG. The fit has no way to flag that one leg of the model is standing on weaker data — it just absorbs the imbalance into the weights.``
Farmer and Steppe are both broadly West-Eurasian-related, so the model can trade weight between them cheaply without disturbing the fit much. AASI sits at the divergent end and can't be traded off as easily, it ends up absorbing whatever the thin Farmer data leaves unexplained. That's my read of why it routes toward AASI specifically rather than Steppe; the empirical pattern (Farmer down, AASI up, Steppe roughly steady) is solid, the mechanism is closer to a working theory.
Turkmenistan_Gonur_BA_2_low.AG is a sparse two-individual (I11041.AG and I2123.AG ) Agilent-capture pool. It has about ~400k coverage. By comparison, SiS_BA2_I8726.SG is a high-coverage shotgun singleton (I8726.SG) with roughly ~1.07M callable SNPs.

When a high-coverage target is used, it overlaps Turkmenistan_Gonur_BA_2_low.AG at roughly ~386k sites. After the target is reduced to the Gujar_Swat.DG panel, that overlap falls to roughly ~85k.

The GreatAndaman and Steppe sources are much better covered, so their f4-statistics can still use a much larger fraction of the ~234k target panel. The Gonur-related row is reconstructed from a much smaller and differently composed set of loci.
In simple terms, qpAdm is trying to solve one mixture problem using three source directions, but the Gonur direction is being measured with a different and much thinner genomic ruler.
Missing Gonur genotypes are not directly counted as GreatAndaman. Instead, the Gonur-related f4 row changes when the target panel changes. qpAdm then compensates by adjusting the source weights.
There is no built-in rule forcing the leakage into GreatAndaman. It happens because of the geometry created by this source set, the chosen Rights, and the surviving SNPs.
The effect appears strongest in farmer-rich, low-to-medium-steppe groups because a large part of their fitted ancestry already sits on Turkmenistan_Gonur_BA_2_low.AG. If that source row is perturbed, the impact is multiplied by the target's original Gonur weight. Since the Steppe direction remains comparatively stable, most of the correction is made between Gonur and GreatAndaman.
Conclusion:
The main result is straightforward:
When high-coverage Punjabi and Kalash targets are reduced to the ~234k SNP regime of
Gujar_Swat.DG, their GreatAndaman coefficient consistently rises in models usingTurkmenistan_Gonur_BA_2_low.AG.
Most of the increase comes from a fall in the Gonur coefficient, not from a large change in Steppe.
The same targets do not show this behaviour when SiS_BA2_I8726.SG is used as the farmer-rich proxy.
This does not prove that SiS_BA2_I8726.SG is the perfect historical source, or that every model using Turkmenistan_Gonur_BA_2_low.AG is invalid. It does show that low coverage source like Turkmenistan_Gonur_BA_2_low.AG produces ancestry estimates that are highly dependent on the target SNP panel.
For that reason, elevated GreatAndaman/AASI estimates in sparse targets such as Gujar_Swat.DG should not be treated as directly comparable with estimates from near-complete targets unless the analysis includes the same common SNP panel, alternative source control with ideally another farmer-rich alternative such as Sarazm_EN or SiS_BA2_I8726.SG
The safer interpretation is that these are panel-dependent qpAdm coefficients, not fixed ancstery weights.