How-To Guide

HarappaWorld Results Explained: GEDmatch Guide

🧬Staring at "S-Indian 48%, Baloch 34%" and wondering what it means for your family? This guide decodes it. When you want a readable state-level breakdown from the same raw file, Helixline starts at $25. Upload your raw data →

For more than a decade, HarappaWorld has been the GEDmatch calculator South Asians reach for. You upload your 23andMe or AncestryDNA file and get 16 percentages with names like S-Indian, Baloch, Caucasian and NE-Euro. Then come the questions. Am I part Baloch? Why do I have Caucasian ancestry? What does a "distance" of 3.2 to "Punjabi" mean?

This guide explains every component in plain English, the patterns South Asian profiles typically produce, how to read Oracle and Oracle-4, how to run the calculator, and what a 2012 tool cannot tell you.

The short version: HarappaWorld is a free ADMIXTURE calculator released in 2012 by the Harappa Ancestry Project. Its 16 components are statistical clusters named after the populations where each one peaked, not real ancestral peoples. For South Asians, S-Indian and Baloch usually dominate, with Caucasian and NE-Euro rising towards the north-west. Oracle turns your percentages into a list of similar reference populations. It is a useful rough guide, not a state-level breakdown.

What is HarappaWorld?

HarappaWorld is an admixture calculator created by Zack Ajmal for the Harappa Ancestry Project, a volunteer project that collected South Asian genotype data from 2011 onwards. According to the announcement post dated 4 May 2012, it uses populations from all over the world, was built from about 188,000 SNPs, and settled on 16 components (K=16) because that gave the lowest cross-validation error. It reached GEDmatch on 21 May 2012. The project has been inactive since its author said in December 2017 that he could not return to it in the near future.

GEDmatch still lists HarappaWorld among its free admixture calculators, alongside others such as Eurogenes K36 and MDLP World-22 (see the GEDmatch Admixture tool).

The one rule for reading HarappaWorld: components are not peoples

ADMIXTURE, the software behind HarappaWorld, finds a set number of statistical clusters in reference data and estimates what share of your genome resembles each. Humans add the names afterwards. The project says so directly: the components "do not necessarily represent real ancestral populations", the names "should be thought of as mnemonics", and they were chosen "based on which populations in my data these components peaked in".

So a "Baloch" percentage does not mean a Baloch great-grandparent. It means part of your genome resembles a cluster that happens to be most concentrated in Baloch-speaking samples in a 2012 reference set. The project also noted that its standard error can be about 1%, so a 1% "exotic" result may just be noise.

Every HarappaWorld component, explained

The "peaks in" column below comes from the Harappa Ancestry Project's own published population averages. The last column describes what the component usually signals in a South Asian result.

Component Peaks in (2012 reference data) What it usually means for South Asians
S-Indian South Indian tribal groups such as the Paniya and Irula The main South Asian component; tends to be higher in South Indian profiles.
Baloch Brahui, Balochi and Makrani samples The second big South Asian component, common everywhere; higher towards the north-west.
Caucasian Georgian, Abkhazian and Armenian samples Usually modest; rises in Punjabi, Kashmiri, Sindhi and Pashtun profiles.
NE-Euro Finnish, Lithuanian and Russian samples Commonly 5-15% in north Indian and Pakistani profiles.
SE-Asian Iban, Dai and Cambodian samples Small in Bengali averages; much higher in some north-eastern groups such as the Garo and Khasi.
Siberian Nganasan, Evenki and Yakut samples Usually 0-3%; higher in some Himalayan profiles.
NE-Asian Japanese samples, and Naga groups of north-east India Higher in north-east Indian and some Nepali profiles.
Papuan Papuan and Melanesian samples 1-4% in some South Indian tribal averages; often noise in an individual.
American Surui, Karitiana and Pima samples Near zero; small values are usually noise.
Beringian Chukchi, Koryak and Greenlandic Inuit samples Near zero; small values are usually noise.
Mediterranean Sardinian and Basque samples Low; small amounts in some north-western averages.
SW-Asian Saudi, Bedouin and Qatari samples Small amounts in some Baloch, Sindhi, Pashtun and Kerala averages.
San San and !Kung samples Near zero.
E-African Gumuz, Anuak and Sudanese samples Near zero unless there is recent East African ancestry.
Pygmy Mbuti and Biaka samples Near zero.
W-African Yoruba and Dogon samples Near zero unless there is recent West African ancestry.

S-Indian

S-Indian is the component most closely associated with the indigenous ancestry of the subcontinent. It peaks at over 80% in South Indian tribal samples. It is loosely comparable to the Ancestral South Indian (ASI) idea from Reich et al., Nature, 2009, but it is not identical: ASI is a modelled ancestral population, while S-Indian is a cluster fitted to 2012 data.

Baloch

Baloch is the component people ask about most. It peaks in Brahui, Balochi and Makrani samples but is widespread, commonly around a third of a South Indian profile and more in north-western ones. Zack Ajmal himself suggested it may be "a mixed component". A high Baloch value is normal for South Asians and does not imply recent ancestry from Balochistan.

Caucasian and NE-Euro

These West Eurasian related components tend to rise towards Kashmir, Punjab, Sindh and the Pashtun regions. Hobbyists often read NE-Euro as partly tracking Steppe-related ancestry; treat that as a loose correlation, not a measurement. Neither means ancestors from the Caucasus or the Baltic.

Typical South Asian patterns (tendencies, not rules)

The project's published averages show some broad tendencies. They come from small, self-selected groups, so treat them as rough orientation.

Individuals vary a lot, and your own result may not look like your region's average.

Want a state-level answer instead of 16 percentages?

Upload the raw file you used on GEDmatch for a plain-English state-level breakdown and community comparisons. From $25, results within 1-2 days.

Upload Your Raw Data

How to read Oracle and Oracle-4 population matches

After the admixture percentages, GEDmatch offers the HarappaWorld Oracle. Oracle compares your percentages with the average percentages of every reference population in the project and ranks them by distance. Lower distance means your percentages look more like that population's average.

Three cautions: if your top matches sit within a small range of distances they are practically tied, so the first name is not "your" population; Oracle can only name groups that were sampled in 2012; and distance is arithmetic on percentages, not shared DNA like centimorgans.

Why a 2012 calculator has limits today

HarappaWorld was built before the major ancient-DNA studies of South Asia. The large 2019 study of ancient genomes from the region (Narasimhan et al., Science, 2019) described South Asian ancestry in terms of Iranian farmer related ancestry, Steppe pastoralist ancestry and AASI (Ancient Ancestral South Indian) ancestry, and how they combined into the ANI and ASI clusters. HarappaWorld's components predate that framework, so there is no neat conversion: Baloch is not "Iranian farmer", NE-Euro is not "Steppe", and S-Indian is not "AASI", even if they overlap loosely. If you want the modern ancient-DNA picture, see our guides to ANI and ASI and Steppe ancestry in India.

Newer GEDmatch calculators, such as the PuntDNAL projects, include components built from ancient samples. Calculators split the same genome differently, so their numbers are not directly comparable.

How to run HarappaWorld on GEDmatch, step by step

  1. Download your raw data from 23andMe, AncestryDNA, MyHeritage, FamilyTreeDNA or LivingDNA, and keep the original .txt or .zip unmodified.
  2. Create a GEDmatch account. The calculators require registration, and uploading your file is required to use them.
  3. Upload your file using the DNA file upload option on your dashboard, and choose a privacy level when asked.
  4. Wait for processing. Once your kit is ready you will get a kit number.
  5. Open the Admixture tool from your dashboard, choose the HarappaWorld project, and enter your kit number.
  6. Run Admixture Proportions, then the Oracle option for population matches. Menus and labels change over time.

A short note on GEDmatch's policies. GEDmatch is a genealogy matching database as well as a toolbox. Under its terms of service, each kit gets a privacy level: Private, Opt-in (includes comparison with kits submitted for law enforcement cases involving violent crime or unidentified remains), Opt-out or Personal Research. Services are not intended for under-16s, and you should only upload DNA you are authorised to share. Choose the setting you are comfortable with.

What HarappaWorld cannot tell you

How Helixline fills the gap

Helixline analyses the same unmodified file you uploaded to GEDmatch. Upload Ancestry ($25 / £19 / CA$34 / ₹1,999, one-time) gives a state-level breakdown, 140+ community comparisons with a closest-match shortlist, ANI/ASI composition, chromosome painting, ancient DNA similarity (probabilistic, for historical context) and Y-DNA and mtDNA haplogroups if your file includes those markers. Upload Complete ($50 / £39 / CA$68 / ₹3,999) adds 30+ wellness traits, carrier screening for conditions common in South Asian populations, pharmacogenomics insights and PDF reports, for information only, not diagnosis.

How it works: choose your plan and currency on helixline.in/upload and pay (Razorpay in 8 currencies, or PayPal in USD). We email you a dashboard link; sign in with the same email and upload your file in about 2 minutes. Results typically arrive within 1-2 days. Full refund until analysis starts or if your file cannot be processed (upload terms). Data is encrypted and stored in India under the DPDP Act 2023, the raw file is auto-deleted within 30 days, and it is never sold.

Frequently Asked Questions

What does S-Indian mean in HarappaWorld?

S-Indian is a statistical component that peaked in South Indian tribal groups such as the Paniya and Irula in the 2012 Harappa reference data. It makes up a large share of almost every South Asian profile and tends to be higher in South Indian results. It is loosely related to Ancestral South Indian (ASI) ancestry but is not the same thing.

Why is my Baloch percentage high?

Because that is normal for South Asians. The Baloch component peaks in Brahui, Balochi and Makrani samples but is widespread across the subcontinent, often a third or more of a South Asian profile and higher towards the north-west. It does not mean recent ancestors from Balochistan; the name only labels where the cluster was most concentrated.

Which GEDmatch calculator is best for South Asians?

There is no single right answer. HarappaWorld remains popular because its reference data included many South Asian volunteers, while newer calculators such as the PuntDNAL projects add components built from ancient samples. Compare patterns rather than exact numbers. For a state-level breakdown or community comparisons, a South Asia focused upload report is more direct.

Is HarappaWorld accurate?

It is a reasonable rough guide to broad patterns, such as how much of your profile leans south or north-west, but it is not precise. It was built in 2012 on about 188,000 SNPs and a volunteer reference set, the component names are mnemonics rather than real ancestral peoples, and the project itself noted a standard error of about 1%. Small percentages and the exact order of Oracle matches should not be over-interpreted.

How is Helixline different from GEDmatch calculators?

GEDmatch calculators return world-scale percentages and nearest reference averages that you interpret yourself. Helixline analyses the same raw file to give a state-level breakdown, similarity to 140+ South Asian community reference groups, ANI/ASI composition, chromosome painting and haplogroups (where your file includes those markers), in plain English. Upload Ancestry is $25 one-time, with results within 1-2 days.

Ready to go beyond 16 percentages? Upload your raw data to Helixline from $25. For more background, read how to interpret 23andMe and AncestryDNA results, how to download and analyse your raw DNA data, or the genetic differences between North and South Indians.

Decoded your HarappaWorld results? Get a readable state-level breakdown from $25 Upload Your Raw DNA