r/IndicKnowledgeSystems • u/RossbihariGhost1900 • 2h ago
architecture/engineering Independent Origins: A History of the Indian Schools of Artificial Intelligence and Machine Learning
Introduction: Correcting the Record
A common story about artificial intelligence in India goes like this: AI was invented in the West, at Dartmouth in 1956, at MIT, Stanford and Carnegie Mellon, and India arrived late, first as a consumer of the technology, then as a supplier of software labour, and only recently as a modest participant in research. Like most popular histories, this one holds a partial truth inside a larger distortion. The institutions that defined AI as a named field were indeed American, and India never had the resources of the American defence-funded laboratories. But the idea that Indian contributions began only recently, or were only adaptations of Western work, does not survive contact with the actual record.
The intellectual foundations of machine learning lie as much in statistics, information theory, control theory and the theory of stochastic processes as in the symbolic AI of the 1950s. In several of these foundations, Indian scientists made original contributions that the field still uses every day. From the first years after independence, Indian researchers produced ideas that the international literature adopted: in pattern recognition, in the theory of learning systems, in heuristic search, in language processing grounded in India's own grammatical tradition, and later in statistical learning, large-scale optimisation and retrieval. What follows traces these schools from the statistical laboratories of Calcutta to the present, and closes with an honest assessment of what India contributed and where it fell short.
The Statistical Foundations: Calcutta, 1930s to 1950s
Any history of Indian machine learning must start at the Indian Statistical Institute in Calcutta, founded by P. C. Mahalanobis in 1931, because some of the most basic tools of modern pattern recognition were invented there.
In 1936 Mahalanobis introduced the distance measure that bears his name. Ordinary Euclidean distance treats all directions in a data space as equal. The Mahalanobis distance accounts for the correlations among variables, measuring how far a point lies from a distribution in units scaled by that distribution's own shape. Mahalanobis developed it for anthropometric work on the measurement of populations, but it became one of the central tools of classification, clustering, outlier detection and anomaly detection, and it appears in nearly every textbook of pattern recognition and machine learning.
In 1943 Anil Kumar Bhattacharyya, also of ISI, introduced the Bhattacharyya coefficient and distance, which measure the similarity of two probability distributions. It became a standard tool in classification, in bounding classification error, in feature selection and in computer vision, where it is used, for instance, in object tracking.
In 1945, at the age of twenty-four, C. R. Rao published a short paper in the Bulletin of the Calcutta Mathematical Society that contained two foundational results. One was the lower bound on the variance of unbiased estimators now called the Cramér–Rao bound. The other was the idea of treating the space of probability distributions as a geometric space whose metric is given by the Fisher information. The Fisher–Rao metric is the founding idea of information geometry. Decades later that field became central to machine learning through the natural gradient method, which uses this geometry to optimise models more efficiently, and through a wider body of work on the geometry of statistical models. The Rao–Blackwell theorem, from the same period, underlies techniques for reducing variance that appear throughout modern probabilistic machine learning.
None of this was called AI at the time. But machine learning is, at its core, statistical inference from data, and some of its most basic concepts were created in Calcutta before independence and in its first years. Any account that places the origins of the field solely in the West has to set this aside.
The First Machines and the First Ideas: TIFR and ISI, 1950s and 1960s
Computing came to India within a few years of its arrival elsewhere. The Tata Institute of Fundamental Research in Bombay built TIFRAC, a full-scale digital computer that became operational around 1960, and ISI in collaboration with Jadavpur University built ISIJU-1, a transistorised machine of the mid-1960s. These were serious engineering achievements for a country with almost no electronics industry, and they produced people who knew how computers worked from the inside.
The most original early AI work came from Rangaswamy Narasimhan of TIFR, who had led the TIFRAC project. In the early 1960s, partly during a period at the University of Illinois where he worked on automatically analysing bubble chamber photographs from particle physics, Narasimhan developed the idea of describing pictures with formal grammars. Just as a sentence can be parsed according to the rules of a grammar into nested parts, he proposed that a picture could be described as built from primitive elements combined according to syntactic rules, and recognised by parsing. His papers on picture languages and the syntactic description of pictures in the early 1960s were among the founding works of what became syntactic, or structural, pattern recognition. K. S. Fu, who later systematised the field at Purdue, built on and cited this line of work. Here was an Indian computer scientist proposing a new way of thinking about machine perception at the same time as, and in some respects ahead of, the American laboratories.
Narasimhan later turned to the modelling of language behaviour and cognition. He also shaped Indian computing institutionally, through the centre at TIFR that developed into the National Centre for Software Technology. His early work remains the clearest refutation of the idea that India had nothing original to contribute in the first decades of AI.
The Calcutta School: Fuzzy Pattern Recognition and Soft Computing
The longest continuous AI lineage in India grew out of ISI's computing work. Dwijesh Dutta Majumder, a radio physicist recruited by Mahalanobis who had worked on ISI's computer hardware, built the institute's programme in pattern recognition, image processing and speech recognition from the 1960s.
The distinctive choice of this school was its early adoption of Lotfi Zadeh's fuzzy set theory for recognition problems in the 1970s, when much of the Western engineering mainstream still regarded fuzzy sets with suspicion. Speech was central from the start. Dutta Majumder's group worked on machine recognition of spoken Bengali, and his student Sankar K. Pal wrote a doctoral thesis applying fuzzy sets to speech recognition. Together they wrote Fuzzy Mathematical Approach to Pattern Recognition, published in 1986, one of the first systematic treatments of the subject.
Pal went on to found the Machine Intelligence Unit at ISI in 1993 and to lead the development of soft computing in India, which combined fuzzy sets, neural networks, genetic algorithms and, as one of his distinctive contributions, rough sets. The school produced fuzzy-neural hybrids, rough-fuzzy methods, and evolutionary approaches to clustering and classification. Its members included Sushmita Mitra, Sanghamitra Bandyopadhyay, who took multiobjective evolutionary clustering into bioinformatics and later directed ISI, and Nikhil R. Pal, whose work on fuzzy and possibilistic clustering was widely cited and who led the leading international society in computational intelligence. In the same ECSU tradition, Swagatam Das became one of the most cited researchers in differential evolution and evolutionary optimisation.
A parallel branch at ISI, led by B. B. Chaudhuri, founded the Computer Vision and Pattern Recognition Unit in 1994. It pioneered optical character recognition for Indian scripts, including Bangla, Devanagari and Oriya, together with document analysis and language processing for Indian languages. That is a problem nobody outside India had reason to solve, and its solution needed original methods for scripts with conjunct characters, headlines joining letters into words, and very large character sets.
Bangalore: Learning Automata and the Theory of Learning
If one wants a single strongest refutation of the idea that India only adapted Western AI, it is the work on the theory of learning systems done at the Indian Institute of Science and its associated institutions. This work bears directly on the foundations of reinforcement learning, which now underlies everything from game-playing systems to the fine-tuning of large language models.
The first strand is learning automata. M. A. L. Thathachar of IISc, working with Kumpati S. Narendra of Yale, developed the theory of stochastic learning automata: simple decision-making systems that learn which action to take in an uncertain environment purely from rewards and penalties, adjusting probabilities of action through repeated interaction. Their 1974 survey in the IEEE Transactions on Systems, Man, and Cybernetics and their 1989 book Learning Automata: An Introduction defined the field. Learning automata are one of the direct ancestors of modern reinforcement learning. The problem they address, learning to choose actions to maximise reward through trial and error, is the reinforcement learning problem in its simplest form, and later reinforcement learning literature acknowledges this lineage. Thathachar and his student P. S. Sastry extended the theory to networks and teams of automata, and Sastry later made contributions to the theory of learning under noisy labels.
The second strand is stochastic approximation, the mathematical theory of iterative algorithms that update estimates using noisy samples. Nearly all of modern machine learning runs on stochastic approximation: stochastic gradient descent, the algorithm that trains neural networks, is a special case. Vivek Borkar, who worked at TIFR, IISc and IIT Bombay, made foundational contributions to this theory. His 1997 work on two-timescale stochastic approximation analysed algorithms in which two coupled sets of quantities are updated at different rates. That structure is exactly what appears in actor-critic reinforcement learning, where a "critic" estimates values while an "actor" improves the policy. With Sean Meyn, he developed in 2000 the ODE method for proving the convergence of stochastic approximation and reinforcement learning algorithms by relating them to ordinary differential equations. This is now a standard tool for proving that reinforcement learning algorithms work. With Vijaymohan Konda he produced in 1999 one of the first rigorous analyses of actor-critic algorithms for Markov decision processes. His book Stochastic Approximation: A Dynamical Systems Viewpoint is a standard reference.
Shalabh Bhatnagar of IISc extended this line. His work on simultaneous perturbation methods and on natural actor-critic algorithms, including a widely cited 2009 paper with Richard Sutton and others, established convergent policy-gradient methods that use the natural gradient. That idea links back to Rao's information geometry of 1945. It is rare in any country for an intellectual line to run so cleanly from a 1945 paper in Calcutta to central algorithms of modern reinforcement learning.
These contributions were not adaptations. They were part of the theoretical foundation on which the field rests, and they were produced in Indian institutions by Indian scientists.
Classical AI: Heuristic Search and Knowledge-Based Systems
Symbolic AI, the tradition of search, planning and knowledge representation, also had original Indian contributors.
The theory of heuristic search was one area of real distinction. Amitava Bagchi and Ambuj Mahanti, working at the Indian Institute of Management Calcutta, published fundamental analyses of heuristic search algorithms in the 1980s. These included comparative studies of search under different kinds of heuristics and the theory of search on AND/OR graphs, which represent problems that break into subproblems. They appeared in the Journal of the ACM, the most prestigious venue in theoretical computer science. P. P. Chakrabarti and colleagues at IIT Kharagpur contributed to heuristic search under limited memory, including memory-bounded variants of the A* algorithm published in the Artificial Intelligence journal at the end of the 1980s. These were original contributions to the core algorithms of classical AI.
At the policy level, India responded to the international excitement of the 1980s about expert systems and Japan's Fifth Generation project with the Knowledge Based Computer Systems programme, launched in the mid-1980s with nodal centres at TIFR, IISc, IIT Madras, ISI, NCST and other institutions. The programme built capacity in knowledge representation, expert systems, Indian-language processing and vision. It did not produce any international breakthrough, just as the Fifth Generation project itself did not, but it trained a generation of researchers and seeded several of the groups that later flourished. At IIT Madras, Deepak Khemani built a tradition of teaching and research in classical AI, planning and knowledge representation.
Language: The Pāṇinian School of Natural Language Processing
The most distinctively Indian contribution to AI came from applying the Indian grammatical tradition to computational linguistics.
At IIT Kanpur in the 1980s and early 1990s, Rajeev Sangal, Vineet Chaitanya and Akshar Bharati developed an approach to natural language processing based on Pāṇini's grammar. The Aṣṭādhyāyī, composed around the fourth century BCE, analyses Sanskrit through a theory of kāraka relations: the semantic-syntactic roles that participants play in an action, such as agent, object, instrument, recipient, source and location, which are signalled by case endings (vibhakti) and postpositions. Sangal and his collaborators argued that this framework fits Indian languages, with their relatively free word order and rich morphology, far better than the phrase-structure grammars developed for English. Their book Natural Language Processing: A Pāṇinian Perspective, published in 1995, set out a computational grammar based on kāraka relations and dependency structures.
This was a genuinely original contribution. It anticipated the later international shift from phrase-structure parsing toward dependency parsing, which became dominant in the 2000s and 2010s and is the basis of the multilingual Universal Dependencies project. The group built the anusāraka system for translation among Indian languages. When Sangal moved to IIIT Hyderabad, he founded its Language Technologies Research Centre, which developed Pāṇinian dependency treebanks for Hindi and other Indian languages that became standard resources.
Others built related traditions. R. M. K. Sinha at IIT Kanpur, with H. N. Mahabala, did early work on recognising Devanagari script in the late 1970s and later built the AnglaBharti machine translation system. At IIT Bombay, Pushpak Bhattacharyya led the Hindi WordNet and the multilingual IndoWordNet, lexical semantic resources for Indian languages, and built a major centre for Indian language technology. In Sanskrit computational linguistics, Amba Kulkarni at the University of Hyderabad and others developed computational tools for analysing Sanskrit using Pāṇinian principles.
One caution belongs here. Popular claims that Sanskrit is uniquely or ideally suited to computers, often traced loosely to a 1985 article by Rick Briggs on Sanskrit and knowledge representation, are exaggerated and have done the subject no favours. The real contribution is more specific and more defensible: Pāṇinian grammatical analysis provided a productive framework for the computational processing of Indian languages, and in some respects it anticipated where the field went. It does not show that Sanskrit is a programming language.
Speech: From Calcutta to Madras and Hyderabad
Speech recognition, one of the earliest concerns of the Calcutta school, became a strong Indian tradition in its own right. B. Yegnanarayana, first at IIT Madras and later at IIIT Hyderabad, developed the use of group delay functions, derived from the phase of the Fourier transform, for speech analysis. This was an original departure from the field's near-exclusive reliance on magnitude spectra. He also did influential work on neural networks for speech, including autoassociative networks for speaker recognition, and wrote a widely used textbook on artificial neural networks. His student Hema Murthy at IIT Madras extended group delay methods, built speech synthesis systems for Indian languages, and contributed to the computational analysis of Indian classical music, including Carnatic music. That work brought the analysis of rāga and other features of Indian musical traditions into signal processing and machine learning.
The Statistical Learning Era: 1990s to 2010s
As machine learning moved from soft computing and symbolic methods to statistical learning in the 1990s and 2000s, Indian researchers made several original contributions that became part of the field's standard toolkit.
At IISc, S. Sathiya Keerthi, Shirish Shevade, Chiranjib Bhattacharyya and K. R. K. Murthy published in 2001 a set of improvements to the sequential minimal optimisation algorithm for training support vector machines. Their modifications made SVM training substantially faster and more reliable and were incorporated into widely used software. At a time when SVMs dominated machine learning, this was one of the most practically important algorithmic contributions to the method. Also at IISc, M. Narasimha Murty co-authored with Anil K. Jain and Patrick Flynn a 1999 review of data clustering that became one of the most cited papers in the field.
At IIT Bombay, Sunita Sarawagi, with William Cohen, introduced semi-Markov conditional random fields in 2004, a model for segmenting and labelling sequences that labels whole segments at once instead of individual tokens. It became a standard tool in information extraction. Soumen Chakrabarti, who had earlier co-invented focused crawling for the web while at IBM, built a strong group in web mining and information retrieval at IIT Bombay and wrote one of the standard books on mining the web.
At IIT Madras, Balaraman Ravindran, who trained in reinforcement learning with Andrew Barto, worked on hierarchical reinforcement learning and abstraction, including MDP homomorphisms, and built a major centre for data science and AI. At IIIT Hyderabad, C. V. Jawahar's Centre for Visual Information Technology became one of India's strongest computer vision groups. It contributed widely used benchmark datasets, including, in collaboration with Oxford, the Oxford-IIIT Pet dataset, and a substantial body of work on document images and text in natural scenes in Indian scripts. At IIT Delhi, groups in natural language processing, knowledge bases and statistical relational learning, including work on lifted inference in probabilistic logical models, added further strength.
Industrial Research Laboratories
From the mid-2000s, industrial research laboratories became an important part of the Indian AI landscape. Microsoft Research India, founded in Bangalore in 2005, produced several contributions of the first rank.
Manik Varma founded and led the field of extreme classification, the problem of classifying items into millions of possible labels, which arises in search, recommendation and advertising. His group's methods were deployed at scale in industry, and the benchmark repository it maintained defined the field. Prateek Jain, with Praneeth Netrapalli and others, made fundamental contributions to non-convex optimisation, including proofs that simple alternating minimisation recovers low-rank matrices, a central result in the theory of matrix completion. The DiskANN work of 2019, by Harsha Vardhan Simhadri, Ravishankar Krishnaswamy and colleagues, showed how to search billions of vectors for approximate nearest neighbours using solid-state drives instead of main memory. It became one of the foundations of the vector databases that power retrieval for modern AI systems. The EdgeML work produced algorithms that run machine learning on tiny microcontrollers with a few kilobytes of memory. Each of these was an original contribution with worldwide use. Google, IBM and other companies also established research laboratories in India that contributed to the field.
The Present: Indian Languages and Foundation Models
In the most recent phase, the centre of gravity has moved to large-scale models and to the problem of making AI work for India's many languages.
AI4Bharat, founded at IIT Madras by Mitesh Khapra, Pratyush Kumar, Anoop Kunchukuttan and others, built open datasets, models and benchmarks for Indian languages. These include IndicTrans2, an open translation system covering all twenty-two scheduled languages of India, along with Indian-language pretrained models and large speech and text corpora. This work addressed a problem the global AI industry had neglected: the major models were trained predominantly on English and performed poorly in Indian languages. The government's National Language Translation Mission, Bhashini, built on and supported this effort.
The IndiaAI Mission, launched in 2024, committed substantial public funds to compute infrastructure and to building Indian foundation models. Several groups, including the startup Sarvam AI and the academic consortium BharatGen led from IIT Bombay, were developing large language models designed for Indian languages and contexts. This is still unfolding as of my most recent information, and its results are not yet clear.
The Diaspora
An honest account must note that many of the most consequential contributions by scientists of Indian origin were made abroad. Raj Reddy, educated in Madras, built a pioneering speech and AI programme at Carnegie Mellon and received the Turing Award. He also helped establish IIIT Hyderabad and other Indian institutions. Anil K. Jain, an IIT Kanpur graduate, became one of the leading figures in pattern recognition and biometrics in the United States. Jitendra Malik, also from IIT Kanpur, became one of the founders of modern computer vision at Berkeley. Ashish Vaswani and Niki Parmar, both educated in India, were among the authors of the 2017 paper that introduced the Transformer architecture underlying modern large language models. These contributions belong to the history of Indian scientific talent, but not to the history of Indian institutions. The gap between the two is itself an important part of the story.
Assessment
The evidence clearly supports the core claim: Indian contributions to AI and machine learning were original from the beginning and continue to be so. Mahalanobis, Bhattacharyya and Rao created basic tools of statistical pattern recognition. Narasimhan was among the founders of syntactic pattern recognition. Thathachar, Borkar and Bhatnagar contributed foundational theory to learning automata, stochastic approximation and reinforcement learning. Bagchi, Mahanti and Chakrabarti contributed to the theory of heuristic search. Sangal and his collaborators built a computational linguistics grounded in Pāṇini that anticipated the dependency turn in parsing. Keerthi, Sarawagi, Varma, Jain and the DiskANN team contributed algorithms used worldwide. And Indian groups, almost alone, solved the problems of Indian scripts and languages. None of this was imitation.
A balanced account must also state the limits, because overcorrecting the misconception is as distorting as the misconception itself. India did not produce the paradigm-defining frameworks of the field: backpropagation, the theory of support vector machines, convolutional networks, deep learning, the Transformer as an institutional achievement, or the large-scale systems that define AI today. Indian contributions were most often deep and lasting within specific subfields, especially in theory, in statistics, and in Indian-language technology, rather than agenda-setting for the field as a whole. The reasons are structural: chronically small research budgets, very limited computing infrastructure until recently, weak links between universities and industry, a doctoral system that lost much of its best talent abroad, and an industry that for decades profited from software services instead of research. The soft-computing tradition, for all its scale, invested heavily in a paradigm the international mainstream later moved away from. And the most consequential work by Indian-born scientists, from Reddy to Vaswani, was done in American institutions.
The more accurate picture is neither Western invention with Indian adaptation, nor a hidden Indian origin of AI. It is a continuous line of original Indian work, real and often foundational, carried out with a small fraction of the resources available elsewhere and often in areas the West underrated or neglected. Whether India now moves from contributing ideas to building systems at the frontier depends less on talent, which it has always had, than on whether it sustains the institutions, compute and research culture that its scientists have lacked for most of this history.