Every time a doctor draws your blood for a routine test, they’re sampling a chemical fingerprint that’s been shaped by thousands of years of human migration, isolation and mixing. Serum proteins like immunoglobulins, haptoglobin and transferrin come in slightly different molecular versions across populations, and these versions are inherited exactly like blood groups. For biological anthropologists, these variations are a goldmine. They map old population boundaries, trace migration routes and reveal which groups mixed with whom long before anyone had access to DNA sequencing. India, with its extraordinary linguistic and ethnic diversity, offers one of the richest datasets for this kind of research anywhere in the world.
Table of Contents
- What serum protein polymorphism actually means
- The GM system: allotypes on antibody heavy chains
- Regional patterns worth knowing
- The KM system: a companion marker on light chains
- Haptoglobin: an iron-recycling protein with a population story
- Why this matters beyond ancestry mapping
- Transferrin: the body’s iron courier
- The TFD marker and tribal populations
- Group Specific Component: the vitamin D transporter
- Why anthropologists read these numbers as history
What serum protein polymorphism actually means
Serum is the clear liquid left after blood clots, and it’s packed with proteins that perform specific jobs, carrying iron, fighting infection, transporting vitamins. A polymorphism simply means a gene exists in more than one common form within a population, and each form (allele) produces a slightly different protein structure. Because these variants are inherited in simple Mendelian fashion and aren’t usually affected by diet, climate or disease, they became some of the first reliable genetic markers used to study human population history, well before genome sequencing existed.
Four systems are especially important for understanding Indian population genetics: the GM and KM systems (both found on antibody molecules), the Haptoglobin (HP) system, and the Transferrin (TF) system. A closely related fifth marker, the Group Specific Component (GC) system, rounds out the picture.
The GM system: allotypes on antibody heavy chains
The GM (Gamma Marker) system refers to inherited structural variants found on the heavy chains of three types of antibodies: IgG1, IgG2 and IgG3. These aren’t variations in what the antibody targets, they’re small differences in the antibody’s own protein scaffold, and they get passed down from parent to child like any other genetic trait.
In Indian populations, four GM haplotypes dominate: GM5, GM1, the combined GM1,5 and GM1,2, with GM5 being the single most common variant nationally. What makes this system anthropologically powerful is how unevenly these haplotypes are distributed. The GM1,5 haplotype is considered a classic marker of Mongoloid ancestry. Research on endogenous groups of West Bengal found that this haplotype reaches high frequencies among Tibeto-Burman-speaking and Mongoloid-affinity communities such as the Rajbanshi, Rabha and Garo, while it is nearly absent among upper-caste Bengali groups like Rarhi Brahmins.
Regional patterns worth knowing
Indo-European and Dravidian speakers, by contrast, tend to carry the GM5 allele at higher frequencies. A study of the Sikh population of Punjab found that their GM/KM haplotype frequencies resembled those reported for other lower-caste Hindu groups of northwestern India, reflecting shared regional ancestry despite religious and cultural differences. Meanwhile, the Himalayan mountain complex, sitting geographically between South Asia, Central Asia and East Asia, shows practically every GM variant present in the subcontinent. This makes sense: the Himalayas have historically functioned less as a hard barrier and more as a corridor, allowing gene flow from multiple directions over long periods.
The KM system: a companion marker on light chains
While GM sits on the heavy chain of immunoglobulins, the KM (Kappa Marker) system involves allotypic variation on the kappa light chain. Three alleles are recognised: KM1, KM1,2 and KM1,3. Across India, the KM1 allele averages around 0.11 in frequency, though local values swing quite a bit depending on the population sampled.
The geographic pattern for KM1 closely tracks the GM1,5 story. Populations of the Himalayan mountain complex, Mongoloid-affinity groups of the eastern Himalaya, and Tibeto-Burman speakers show elevated KM1 frequencies. Indo-European and Dravidian speakers, along with Scheduled Tribes of the southern peninsula, show markedly lower values. This parallel isn’t a coincidence, since both GM and KM are inherited on immunoglobulin molecules and tend to reflect the same broad ancestral divide between populations with East and Southeast Asian affinities and those with West Eurasian or indigenous South Asian ancestry.
Haptoglobin: an iron-recycling protein with a population story
Haptoglobin is a glycoprotein made of two alpha and two beta chains, and its job is to mop up free haemoglobin released into the bloodstream when red blood cells break down. Without haptoglobin, that free haemoglobin would cause oxidative damage and put unnecessary strain on the kidneys. The protein’s hemoglobin-binding affinity is exceptionally strong, among the tightest protein-protein interactions known in the body.
The polymorphism itself sits almost entirely in the alpha chain, giving rise to the HP1 and HP2 alleles. Across India, HP1 frequency swings dramatically, from close to zero in some groups to as high as 0.40 in others. The Himalayan mountain complex, Mongoloid-affinity populations, and Mon-Khmer-speaking Austro-Asiatic groups tend to show high HP1 frequencies. On the other end of the spectrum, Scheduled Tribes of the Brahmaputra plains, Munda-speaking Austro-Asiatic groups and many southern Indian populations show notably low HP1 frequencies.
Why this matters beyond ancestry mapping
Haptoglobin type isn’t just a historical marker, it has real physiological relevance. Because HP1 and HP2 forms differ in how efficiently they clear haemoglobin, researchers have studied haptoglobin type in relation to malaria resistance, tuberculosis progression and cardiovascular risk. A study among tribal groups of Andhra Pradesh, for instance, documented an HP1 frequency range of roughly 0.056 to 0.159 across five communities, alongside notable variation in the related GC system within the same populations.
Transferrin: the body’s iron courier
Transferrin is a glycoprotein manufactured in the liver whose entire job is ferrying iron through the bloodstream, from the gut and spleen to wherever new red blood cells are being made. Its levels rise when the body senses iron deficiency, since more transporter capacity is needed when iron is scarce, which is part of why transferrin is measured in anaemia workups even today.
Genetically, three main transferrin phenotypes matter for Indian population studies: TFC (with variants C1 through C13), TFD (including D and Chi) and TFB. TFC is overwhelmingly the most common phenotype nationally, with a frequency near 0.991, while TFD sits around 0.008 and TFB around 0.001. But the sub-variants within TFC tell a more textured regional story. Western Indian populations show particularly high TFC1 frequency, while TFC3, generally treated as a marker of European genetic input, appears specifically among north Indian populations. A broader review of transferrin subtypes across the country similarly found a west-to-east frequency cline in the TFC2 allele, rising steadily toward eastern India, hinting at a gradient of admixture rather than a sharp boundary.
The TFD marker and tribal populations
TFD, the slower-moving variant, shows up disproportionately in Scheduled Caste and Scheduled Tribe populations, particularly those living in tropical savannah climate zones, and among Munda-speaking Austro-Asiatic groups. A study of tribal populations in Andhra Pradesh found that a related slow transferrin marker, D-Chi, reached its highest recorded frequency among Koya Dora and Naikpod communities, while being essentially absent among the neighbouring Pardhan group, despite geographic proximity. This kind of fine-grained variation between neighbouring tribes is common in serum protein studies and reflects centuries of relative reproductive isolation, even among communities living within a few kilometres of each other.
Group Specific Component: the vitamin D transporter
The GC system, sometimes treated alongside the core serum protein markers, involves the alpha-2 globulin responsible for transporting vitamin D metabolites through the blood. Its phenotypes include GC 1S, GC 1S1F, GC 1F, GC 2-1S, GC 2-1F and GC 2, arising from two main alleles, GC1 (split further into 1F and 1S subtypes) and GC2.
Nationally, GC1 frequency averages around 0.74, ranging from 0.59 to 0.91 depending on the population. Munda-speaking Austro-Asiatic groups and populations from Andhra Pradesh show particularly high GC1 frequencies. Within the GC1 allele, the two subtypes split along a striking geographic and ancestral line. GC1S, reaching frequencies around 0.49, predominates in Mongoloid-affinity populations of the western Himalaya and among the Onge of the Andaman Islands. A broader Asian-Pacific survey confirmed this pattern at scale, finding that Indian populations show the highest GC1S frequencies anywhere in the region, a pattern the researchers linked to shared genetic affinities with European populations. GC1F, by contrast, appears more often in eastern Himalayan populations, at roughly 0.25 frequency, aligning with the East and Southeast Asian genetic influence that shows up repeatedly across the other serum protein systems in this region.
Why anthropologists read these numbers as history
None of these frequencies exist in isolation. Put together, GM, KM, HP, TF and GC data consistently point toward the same broad genetic geography of India: a west-to-east and south-to-north cline running from populations with West Eurasian affinities (Indo-European and Dravidian speakers) toward populations with stronger East and Southeast Asian affinities (Tibeto-Burman speakers and Mongoloid-affinity groups of the Himalaya and northeast). This pattern has since been confirmed using modern genome-wide data, which describes India’s population structure in terms of four major ancestral components tied to these same linguistic families. What’s remarkable is that anthropologists were already detecting this structure decades earlier using nothing more than blood serum and simple biochemical typing techniques. The Himalayan mountain complex’s unusually mixed marker profile, carrying variants from every major system, also fits this picture neatly. Rather than acting as a wall separating South Asia from Central and East Asia, the Himalayan belt functioned as a zone of contact, absorbing gene flow from multiple directions over millennia.
These markers also carry a caste and tribe dimension that’s distinct from pure geography. Scheduled Tribe and Scheduled Caste populations in several regions show marker frequencies (low HP1, presence of TFD, distinct GM haplotypes) that set them apart from neighbouring caste Hindu populations even when they share the same territory. This reflects patterns of social endogamy, marriage largely confined within one’s own community, that have shaped India’s genetic landscape as much as geography has.
What do you think? If a tribal community and a neighbouring caste group have lived in the same district for centuries, what social or historical factors might explain why their serum protein frequencies still look so different today? And given that these markers were discovered using techniques far simpler than modern DNA sequencing, what does that tell you about how much history can be reconstructed from even basic biological data?
References
- https://pubmed.ncbi.nlm.nih.gov/3111344/
- https://onlinelibrary.wiley.com/doi/abs/10.1002/ajpa.1330750105
- https://pmc.ncbi.nlm.nih.gov/articles/PMC8365284/
- https://pubmed.ncbi.nlm.nih.gov/6153377/
- https://www.ncbi.nlm.nih.gov/books/NBK532928/
- https://pubmed.ncbi.nlm.nih.gov/9686483/
- https://link.springer.com/article/10.1007/bf00291395
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3514343/
Leave a Reply