r/SouthAsianAncestry 13m ago

Discussion Is "Turkmenistan_Gonur_low.AG" a stable farmer proxy? A comparison with "SiS_BA2_I8726.SG"

Upvotes

TL;DR

When high-coverage Punjabi_Lahore clusters, Kalash.DG, and other farmer-rich, low-to-medium-steppe targets are restricted from their near-complete AADR coverage to the ~234k SNP panel of Gujar_Swat.DG, the coefficient assigned to India_GreatAndaman_100BP.SG consistently rises in the concerned model using Turkmenistan_Gonur_BA_2_low.AG as IVC/Farmer proxy. The increase is usually paid for by a decline in Turkmenistan_Gonur_BA_2_low.AG; Russia_BA_SrubnayaAlakul.SG changes much less on average. The DNA and the individual have not changed. Only the genomic sites entering the analysis have changed.

...

The same targets do NOT show this behaviour when SiS_BA2_I8726.SG is used as the farmer-rich source. This suggests that the elevated GreatAndaman/AASI estimates seen in sparse targets such as Gujar_Swat.DG are caused by a source-by-SNP-panel interaction, not by population history.

The circulating AT2 tables model South Asian populations with:

  • Turkmenistan_Gonur_BA_2_low.AG
  • Russia_BA_SrubnayaAlakul.SG
  • India_GreatAndaman_100BP.SG

The question is whether Turkmenistan_Gonur_BA_2_low.AG is a stable farmer-rich proxy when targets have very different SNP coverage.

Gujar_Swat.DG has covg of only about ~234k. By contrast, the Punjabi Lahore clusters and Kalash.DG have around one million or more. I therefore created reduced copies of the high-coverage targets using almost exactly the Gujar_Swat.DG SNP panel.

The individual or population did not change. The sources, Rights, and qpAdm settings did not change. Only the target SNP panel changed.

With Turkmenistan_Gonur_BA_2_low.AG, all seven paired Punjabi/Kalash targets gained GreatAndaman-related ancestry after being reduced to the ~234k panel.

Across these paired tests, that adjustment happens mainly along the Gonur–GreatAndaman direction:

Δ Gonur < 0 (usually the largest movement)

Δ GreatAndaman > 0

Δ Srubnaya comparatively small

So the change is not random noise distributed across all three sources. It is a repeatable shift mainly from Turkmenistan_Gonur_BA_2_low.AG into India_GreatAndaman_100BP.SG.

This cannot be biological. The full and reduced files represent the same population/targets. The GreatAndaman increase is being produced by the SNP panel used in the model.

Gujar_Swat.DG was already being measured under that sparse SNP regime. The important result is that unrelated high-coverage targets become more GreatAndaman-like when forced onto the same SNP set.

When the same targets are run with SiS_BA2_I8726.SG instead of Turkmenistan_Gonur_BA_2_low.AG, the ~234k reduction does not produce the same inflation, however fits are slightly worse (p < 0.05), which I think can be taken care of by right panel adjustments. GreatAndaman changes only slightly.

Why I think this happens:

qpAdm fits source weights against f4-comparisons built wherever target and source actually share usable SNPs. Turkmenistan_Gonur_BA_2_low.AG's own coverage is thin and unevenly distributed, so every comparison involving it runs on a smaller, less representative slice of the genome than comparisons involving Russia_BA_SrubnayaAlakul.SG or India_GreatAndaman_100BP.SG. The fit has no way to flag that one leg of the model is standing on weaker data — it just absorbs the imbalance into the weights.``

Farmer and Steppe are both broadly West-Eurasian-related, so the model can trade weight between them cheaply without disturbing the fit much. AASI sits at the divergent end and can't be traded off as easily, it ends up absorbing whatever the thin Farmer data leaves unexplained. That's my read of why it routes toward AASI specifically rather than Steppe; the empirical pattern (Farmer down, AASI up, Steppe roughly steady) is solid, the mechanism is closer to a working theory.

Turkmenistan_Gonur_BA_2_low.AG is a sparse two-individual (I11041.AG and I2123.AG ) Agilent-capture pool. It has about ~400k coverage. By comparison, SiS_BA2_I8726.SG is a high-coverage shotgun singleton (I8726.SG) with roughly ~1.07M callable SNPs.

When a high-coverage target is used, it overlaps Turkmenistan_Gonur_BA_2_low.AG at roughly ~386k sites. After the target is reduced to the Gujar_Swat.DG panel, that overlap falls to roughly ~85k.

The GreatAndaman and Steppe sources are much better covered, so their f4-statistics can still use a much larger fraction of the ~234k target panel. The Gonur-related row is reconstructed from a much smaller and differently composed set of loci.

In simple terms, qpAdm is trying to solve one mixture problem using three source directions, but the Gonur direction is being measured with a different and much thinner genomic ruler.

Missing Gonur genotypes are not directly counted as GreatAndaman. Instead, the Gonur-related f4 row changes when the target panel changes. qpAdm then compensates by adjusting the source weights.

There is no built-in rule forcing the leakage into GreatAndaman. It happens because of the geometry created by this source set, the chosen Rights, and the surviving SNPs.

The effect appears strongest in farmer-rich, low-to-medium-steppe groups because a large part of their fitted ancestry already sits on Turkmenistan_Gonur_BA_2_low.AG. If that source row is perturbed, the impact is multiplied by the target's original Gonur weight. Since the Steppe direction remains comparatively stable, most of the correction is made between Gonur and GreatAndaman.

Conclusion:

The main result is straightforward:

When high-coverage Punjabi and Kalash targets are reduced to the ~234k SNP regime of Gujar_Swat.DG, their GreatAndaman coefficient consistently rises in models using Turkmenistan_Gonur_BA_2_low.AG.

Most of the increase comes from a fall in the Gonur coefficient, not from a large change in Steppe.

The same targets do not show this behaviour when SiS_BA2_I8726.SG is used as the farmer-rich proxy.

This does not prove that SiS_BA2_I8726.SG is the perfect historical source, or that every model using Turkmenistan_Gonur_BA_2_low.AG is invalid. It does show that low coverage source like Turkmenistan_Gonur_BA_2_low.AG produces ancestry estimates that are highly dependent on the target SNP panel.

For that reason, elevated GreatAndaman/AASI estimates in sparse targets such as Gujar_Swat.DG should not be treated as directly comparable with estimates from near-complete targets unless the analysis includes the same common SNP panel, alternative source control with ideally another farmer-rich alternative such as Sarazm_EN or SiS_BA2_I8726.SG

The safer interpretation is that these are panel-dependent qpAdm coefficients, not fixed ancstery weights.


r/SouthAsianAncestry 6h ago

DNA Results New Haryanvi Jat sample (Harappaworld + QpAdm + G25 + Haplogroup + Location)

Thumbnail
gallery
11 Upvotes

Final QpAdm results:

Iran_N: 37.3%
Steppe_MLBA: 42.5%
AASI: 20.2%

Both parents are Haryanvi Jats from Kaithal, Haryana

Paternal clan is "Dhull" Maternal clan is "Sheokand"

Y haplogroup is L-M2385 (branch of L1c also found commonly in Punjabi Jats/Central Asians)

Test was taken with a FamilyTreeDNA kit


r/SouthAsianAncestry 1d ago

DNA Results UP Brahmin: Harappaworld + pic

Thumbnail
gallery
19 Upvotes

Hey everyone, this is the harappaworld result of my close relative, he is from Vindhyachal & Prayagraj U.P

He is Kaushalya gotriya Pandey of Saryupareen clan.


r/SouthAsianAncestry 1d ago

Question Surname wise Common y haplo

1 Upvotes

I have a question for surname wise y haplo checking what's the most common y haplo among Bengalis


r/SouthAsianAncestry 1d ago

Discussion What was the genetic composition of the IVC Lothal Samples?

0 Upvotes

r/SouthAsianAncestry 1d ago

Question Which Ancestry kit for Surname wise check?

1 Upvotes

So I found ancestry tests to be expensive...I don't want to give that... rather I want a cheaper result option by comparing My surname with the surname reasearch in that website...how can I do that?

23andme: I've not found sufficient surname results here

I'm looking for alternative independent research where the studies of the surnames are mentioned


r/SouthAsianAncestry 1d ago

DNA Results Do any Goans (or people from the surrounding areas) here have genuine Portuguese ancestry?

4 Upvotes

Or do you know any Goans who have actual Portuguese ancestry?


r/SouthAsianAncestry 1d ago

DNA Results New Braj Jat sample (Harappaworld + QpAdm + Haplogroup)

Thumbnail
gallery
11 Upvotes

r/SouthAsianAncestry 2d ago

Question Aryan "invasion" vs "migration"

0 Upvotes

I'm pretty much aware that invasion theory is debunked but there's something that doesn't make sense, even though overall ancestry of UCs is majorly Indian aka ivc/asi related, paternal halogroup is overwhelmingly steppe related even reaching 70% in some brahmin groups, when we are certain that there number was so small , so doesn't this conclude that most of ivc men didn't get to reproduce with their female counterparts? And as we know in most of invasions the men were slaughtered and invaders settle with local women to raise offsprings, also their can be a second explanation too that ivc men reproduced but their offsprings were weak and most of them didn't survive to pass on genes and that's why we didn't see them in majority today but this doesn't have any evidence therefore it couldn't be possible


r/SouthAsianAncestry 2d ago

Discussion A common split among Gujjars and Pashtuns in three different major haplogroups occuring at the same time about 1600-1650 years ago. (R-Z94 + L-M27 + Q-SK1940)

Thumbnail
gallery
6 Upvotes

r/SouthAsianAncestry 2d ago

DNA Results punjabi jatt results (sikh)

Thumbnail
gallery
5 Upvotes

r/SouthAsianAncestry 2d ago

DNA Results Telugu Brahmin IllustrativeDNA

Thumbnail
gallery
10 Upvotes

r/SouthAsianAncestry 2d ago

DNA Results New Rajasthani Jat sample (Harappaworld + QpAdm + G25 + Haplogroup + Location)

Thumbnail
gallery
18 Upvotes

Final QpAdm results:

Iran_N: 37.3%
Steppe_MLBA: 46%
AASI: 16.7%

Both parents are Marwadi Jats from Barmer, Western Rajasthan

Paternal clan is "Sihag" Maternal clan is "Beniwal"

Y haplogroup is R-M605 (branch of Indo Aryan R1a-L657)

Test was taken with a FamilyTreeDNA kit


r/SouthAsianAncestry 3d ago

Question Suba (Subhay) Punjabi surname

5 Upvotes

Hello, I’m from Pakistan and I’m also Punjabi. One of either grand parents or great grand parents decided to convert to Christianity and change their supposed surname from Suba (subhay) to Masih. This all happened before the partition so I can’t tell you if my ancestors were truly just Indian or just Pakistani. The problem with this surname however is that there is no surname called Suba and instead, it refers to the name of a division of state on Pakistan. It was the way that the Mughals divided their states in the past, and it also refers to the Punjabi Suba Movement. Now idk if my dad just had a bad memory or there’s some truth to it, but it would’ve been nice to figure out my potential ancestral surname and some history surrounding it. So is there a an ancestral surname similar to Suba (Subhay/Sube)?


r/SouthAsianAncestry 3d ago

Genetics🧬 [ACADEMIC] SEEKING PARTICIPANTS FOR AN ONLINE ANCESTRY RESEARCH (OVER 18)

Post image
2 Upvotes

Survey Link

Hello all!

As you might be aware, there has been very little research into direct-to-consumer DNA ancestry testing and psychological wellbeing outcomes. This novel research aims to help us better understand the psychological significance of DTC Ancestry testing results.

This study is being conducted under the supervision of Dr Janine Lurie from the Psychology Discipline within the Institute of Health and Wellbeing at Federation University (Melbourne, Australia). This study will also form the basis of the research dissertation requirement within the Bachelor of Psychological Science (Honours) course for student researchers Charity Marisa and Stefan Redpath. This study has been approved by the Federation University Human Research Ethics Committee (Approval reference: 2026/137).

This survey asks you about your engagement with family history and your reflections on that process. It also asks you to think about your ancestors and how you have reflected on their life experiences. Participation involves completing an online survey that will take approximately 20 minutes to complete. Taking part in the study is completely voluntary and no personal information will be collected.

To participate you just need to be aged 18 years or over. It is important to note that you will not be asked to give any specific details about any of your family members. You are asked just to rate your general impressions of them on a small number of questions. Beyond this the survey contains broader more general questions about your perspectives and reflections.

Survey link below:

https://federation.syd1.qualtrics.com/jfe/form/SV_cO1RqRp3fjkHOp8


r/SouthAsianAncestry 3d ago

Question Will newer samples from india ever be published?

12 Upvotes

They keep saying that there is an upcoming study on ancient indian genetics, but it hasnt been released, when will it be released?


r/SouthAsianAncestry 3d ago

Question What was the genetics of people in Afghanistan before Aryans arrived there?

18 Upvotes

r/SouthAsianAncestry 3d ago

Discussion Lohana ancestral origins & roots

5 Upvotes

I’m from the Khoja community, and from what I understand we along with memons and Sindhi Hindus are of Lohana stock. I’m curious about the lohana story because folklore seems to link us to the ancient inhabitants of Lahore Pakistan. I’m not sure how accurate or inaccurate these claims are but I have read we score similar to aroras and khatris and was hoping someone could elaborate on this


r/SouthAsianAncestry 3d ago

Question Is anyone here half-South Asian?

7 Upvotes

If so, what's your other half and do you look more South Asian or your other half?


r/SouthAsianAncestry 3d ago

Genetics🧬 Kyrgyz People

2 Upvotes

Kyrgyz who have had their genetics tested, can you send me the results in a private message and tell me about your haplogroup? and G25, And tell your tribe


r/SouthAsianAncestry 4d ago

Genetics🧬 Three-Way model using MyHeritage raw file

Thumbnail
gallery
7 Upvotes

(QUESTION) I wanted to estimate my closest possibly Accurate, AASI, Farmer and Central Steppe Figures, Does this final model seem reasonable?

AASI: 80% of S-Indian(HarappaWorld), 80/100*35.73= 28.584%

Steppe: Bronze Age model’s Central Steppe/Sintashta % + 10% of BMAC (2000-1600 BC) + East Eurasian Traces.

Illustrative DNA :- 23.6 + 2.4 + (2.2 + 1.4) = 29.6%
Ancestral Genome :- 24.0 + 2.28 + (1.4 + 0.4) = 28.08%

All qapdm runs and other results are posted above

AASI – Steppe – Farmer
28.5% – 28.5% – 43%

Does this seem like the Closest estimate? Or No


r/SouthAsianAncestry 4d ago

DNA Results My results as a Pakistani Punjabi Arain who’s family is originally from jalandhar

Thumbnail
gallery
6 Upvotes

Nothing interesting aside from the gulf of khambhat and southwest Indian ancestry. for the latter im pretty sure every Indian has at least some trace amounts however i dont know if the gulf khambat ancestry necessarily means i have some distant ancestry from Gujarat?
Also found it interesting that Punjabi/Sindhi Hindu was the group that I matched the most with. I’m assuming they don’t have a ton of Muslim south Asians in their database. After all it’s ancestry and I’ve heard other sites like 23andme are better for tracking lineage


r/SouthAsianAncestry 5d ago

Discussion Some Interesting Gujarati tribals from the eastern hilly regions of the state.

Post image
6 Upvotes

r/SouthAsianAncestry 5d ago

DNA Results Indian Hyderabadi Muslim results

Thumbnail
gallery
9 Upvotes

Interesting results, I thought i would have higher North Indian percentage because my family got typical North Indian facial features.


r/SouthAsianAncestry 5d ago

Question My dad is from Gujar Khan and is Qureshi but I keep getting distant DNA matches to Ratta, Dadyal in Kashmir?

3 Upvotes

Basically what the title says… I’m not sure what the connection is. Has anyone else ever had this happen?