MIDF Update II: Building Bharat’s Manuscript and Inscription Infrastructure
By MIDF
The Manuscripts and Inscriptions Digitisation Foundation has moved from concept to institution. The first version of the MIDF repository is operational: records are structured, metadata descriptions are visible, partner workstreams are active, and the next phase of the platform is in development.
This second public update serves a single purpose , to report, with clarity and evidence, what has been built, what is being built next, and why this work is now ready to scale.
Where things stand
The MIDF repository currently holds 2,107 manuscripts and 698 inscriptions, organised across more than 11 languages. Against the full scale of what India holds, these are opening numbers, and they are understood as such.
What they represent, more than any count, is a proof of method. Behind each record is structured metadata. Every entry is searchable, describable, and connected to a system designed to absorb far larger volumes without losing coherence.
This is what has been missing from the field: not digitisation, but a metadata-first architecture that makes digitisation usable.

Image 1: MIDF Repository at a Glance. Source: MIDF Update II.
The problem MIDF is solving
India is believed to possess one of the largest collections of manuscripts and inscriptions in the world. MIDF works with an estimate of around 10 million manuscripts and over 3 lakh inscriptions, though we believe the true number is even higher. Written in more than a hundred scripts and spanning nearly five millennia, these records preserve an extraordinary legacy of religious, scientific, literary, legal, philosophical, and administrative thought.
Alongside them sits an inscriptional record of more than three lakh inscriptions on stone, copper plates, and temple walls that together constitute the documentary foundation of Indian civilisation.
Digitisation has been underway for years, and the work done by government institutions, universities, temples, mutts, libraries, private custodians, and international repositories deserves recognition. A substantial body of material has been scanned, photographed, and partially catalogued, often under considerable resource constraints.
The problem, however, is not the volume of digitisation. It is the absence of common infrastructure beneath it. Records are distributed across platforms, PDFs, hard drives, and local archives with no shared standard connecting them. Many exist only as static images. Some are catalogued. Some are partially searchable. A great many are neither. A researcher today faces structurally the same discovery problem a researcher faced thirty years ago: India’s primary sources are widely distributed, inconsistently described, and weakly connected to one another.
The bottleneck has shifted, but it has not disappeared. Preservation itself remains a major issue in need of attention, as countless manuscripts, documents, recordings, and cultural materials around the world still face risks from physical deterioration, conflict, neglect, and environmental damage.
At the same time, the challenge is no longer only preservation. It is also discovery, interoperability, Unicode conversion, metadata quality, full-text access, transcription, transliteration, and long-term digital durability. A scan preserves the image; it does not, by itself, preserve access.

Image 2: From Scattered Files to Usable Knowledge. Source: MIDF Update II.
Ensuring that preserved materials can be found, read, searched, connected, reused, and sustained across generations is the second half of the work—and that is the work MIDF exists to do.
The infrastructure gap
Before building, MIDF conducted a systematic review of 23 manuscript repository platforms , Indian and international, governmental and academic, institutional and community-led , scored across infrastructure, access, search, and scholarly capability.
The findings were instructive. The average platform scored 15.5 out of 50; the median was 13. Indian-origin platforms averaged 9.9. Only ten of the twenty-three support IIIF, the international standard for image interoperability, and only one of those is Indian-origin. Twelve platforms offer no programmatic access whatsoever. Across all twenty-three, no platform was found with production-grade OCR for Indic manuscript scripts, cross-Indic-script search, or collaborative Indic transcription capability.
The strongest international systems, such as Gallica and the Buddhist Digital Resource Center, are genuine achievements in digital library infrastructure. They offer IIIF, APIs, OCR pipelines, linked data, and advanced search. But they were built for their own collections and their own philological questions. They were not designed for the complexities of India’s textual heritage: dozens of scripts, multiple language families, cross-script search, transliteration, and the integration of manuscripts and inscriptions at national scale.
The field divides cleanly: content depth on one side, infrastructure depth on the other, and the specific problem of Indic-script intelligence sitting unsolved in the gap between them. MIDF is positioned to close that gap.

Image 3: Illustrative Platform Landscape. Source: MIDF Update II.
What MIDF is building
MIDF is building the missing infrastructure layer for India’s manuscript and inscription record. The platform is organised across five interdependent layers.
MIDF is building the missing infrastructure layer for India’s manuscript and inscription record. The platform is organised across five interdependent layers:
This is not a PDF library or a visual archive. The long-term goal is a public knowledge infrastructure , one where manuscripts and inscriptions can be preserved, searched, cited, compared, improved, and connected across institutions and generations.
The product intuition, stated plainly, is close to a Wikipedia for manuscripts and a GitHub for scholarly heritage records: open, collaborative, versioned, and citable. The implementation, however, must be more rigorous than any analogy suggests. Scholarly review, provenance, attribution, custodial rights, institutional agreements, and quality control sit at the centre of the design.

Image 4: MIDF’s Five-Layer Infrastructure Model. Source: MIDF Update II.
Progress to date
MIDF‘s execution has run along four tracks over this period.
Repository and Metadata
The first version of the platform is active, with 2,107 manuscripts and 175 inscriptions now organised and visible. The architecture is metadata-first by design, because metadata is the bridge between preservation and research: without it, a scan is merely an image; with it, the same image becomes searchable, citable, comparable, and eventually machine-readable.
The current build makes structured descriptions publicly visible; the next version will substantially improve search, record architecture, and public usability.
Manuscript Preservation and Partner Digitisation
MIDF has actively supported several key open-access digitisation initiatives to expand preservation efforts across regions:
Separately, unpublished manuscripts are being systematically connected to university research and dissertation pathways, building a reliable channel through which the repository feeds active academic scholarship directly.
The strategic orientation is important: MIDF is not seeking to centralise ownership of collections. The direction is to support digitisation at source, raise standards across the field, and build a common repository architecture in which collections become discoverable without displacing the custodianship of the institutions and families who have preserved them.
Inscriptions and Epigraphy
MIDF treats inscriptions as a primary category of the civilisational record, not a supplementary one. In addition to the 175 inscriptions organised within the MIDF repository, partner Pratnakirti in Dharwad has assembled a registry of over 1,000 inscriptions and more than 71,000 metadata records.
A twelve-to-fifteen-month project to digitise the Epigraphia Indica and Epigraphia Carnatica volumes into Unicode has been launched, alongside a targeted push to document 500 inscriptions that are known to scholars but remain unpublished.
The case for prioritising inscriptions is not only historical. Inscriptions are primary evidence in the strictest sense: they record grants, donations, dynastic succession, temple administration, astronomy, law, social history, geography, and political authority, often with dates and proper names that no other source preserves. Yet much of India’s inscriptional record remains locked in scanned volumes, printed corpora, and specialist files. Without Unicode text, full-text search, and standardised references, the material is effectively inaccessible to the research it could serve. The objective is unambiguous: inscriptions must become searchable records, not only photographed stone.

Image 5: From Inscription Source to Searchable Record. Source: MIDF Update II.
Institutional and Academic Partnerships
MIDF‘s current network includes alignment with the IKS Division of the Ministry of Education; active conversations with Sri Sathya Sai University for Human Excellence, Central Sanskrit University, Chanakya University, and Sri Venkateswara University in Tirupati; and operational partnerships with KVSRI Chennai, Sarasvata Samsthanam (conducting survey work across Southeast Asia, including Indonesia and Sri Lanka), Pratnakirti Dharwad, eGangotri, and eSahitya.
The long-term academic model is to embed unpublished manuscript research into university programmes, so that students and scholars can select, study, transcribe, and publish research on unpublished material as part of structured academic pathways. This turns the repository into a research engine rather than only a storage layer , one that generates new scholarship as a natural consequence of its operation.
Leadership and advisory structure
Manuscript infrastructure of this complexity cannot be built by any single discipline alone. MIDF‘s leadership and advisory base spans scholarship, technology, institution-building, philanthropy, and cultural preservation.
MIDF is led by:
The advisory circle includes Padma Shri Chamu Krishna Shastry, co-founder of Samskrita Bharati and Chairman of the Bharatiya Bhasha Samiti; Padma Shri Dr M.D. Srinivas, Chairman of the Centre for Policy Studies; Arun Yogiraj, fifth-generation sculptor from Mysuru and creator of the Ram Lalla idol in Ayodhya; and Padma Shri Dr J.K. Bajaj, who with Dr Srinivas is co-leading the IKS Library System classification effort , the intellectual framework by which this material will be organised at scale.
The national context: Gyan Bharatam Mission
The Government of India’s Gyan Bharatam Mission has established manuscript preservation as a national priority, with stated objectives that include digitising and cataloguing over one crore manuscripts, creating a National Digital Repository, and integrating AI, OCR, and provenance technologies for smarter access and transcription. This is the appropriate national ambition, and MIDF‘s role is conceived as complementary to it.
What MIDF brings to that effort is execution infrastructure: repository architecture, Unicode-first workflows, partner onboarding, metadata quality standards, and AI-ready data preparation.MIDF has formally sought consideration as an execution support partner for the Mission, including the opportunity to present its infrastructure and roadmap to the concerned agencies.
The operating theory is clear. India’s manuscript and inscription record will require government ambition, private execution capacity, scholarly rigour, donor capital, and institutional standards working in sustained coordination. The scale of the task makes any other arrangement insufficient.
Three-year targets
MIDF‘s next-phase goals are set to be measurable and reported against.

Image 6: Roadmap to 2028. Source: MIDF Update II.
Funding structure and the Phase 1 ask
MIDF is building a ₹15 crore strategic corpus for the 2028 mandate, structured along a two-year roadmap: ₹8 crore allocated to technology and ₹7 crore to strategic grants to the field.
The Phase 1 ask is ₹5 crore, applied across four areas: platform development , the Unicode-compliant architecture, design, and build of the repository; hosting and computing infrastructure to store and serve high-resolution scans reliably; AI-ready data onboarding to structure and tag content from select partner projects; and seed grants to digitisation partners willing to adopt MIDF standards, so that quality propagates across the field.
This is, properly understood, capital expenditure on knowledge infrastructure rather than a donation to a preservation effort. The costs that make the work possible , scanning, metadata labour, server infrastructure, storage, Unicode conversion, software development, scholar review, field coordination, partner grants, and long-term maintenance , are not glamorous. But they are precisely what determines whether this material survives in usable form or merely in pixels. They are budgeted for honestly, and they are what this ask is designed to fund.

Image 7: Funding Structure and Phase 1 Ask. Source: MIDF Update II.
Why this must be built now
The urgency of this work manifests in three ways.

Image 8: Three Forms of Urgency. Source: MIDF Update II.
The first is physical. Manuscripts and inscriptions continue to decay, fragment, and disappear , to moisture, fire, mishandling, and neglect. The pace of loss is not uniform, but it is ongoing, and it is irreversible.
The second is scholarly. Research across multiple disciplines continues to be built on incomplete access. Scholars spend time locating sources rather than studying them, and a substantial body of primary material remains unknown, uncited, and unavailable to the field that would use it.
The third is technological, and increasingly consequential. AI systems learn from structured, machine-readable data. If India’s manuscripts and inscriptions remain as unstructured images, they will not participate in the next generation of knowledge systems. In a research environment increasingly shaped by large-scale AI, material that cannot be searched, read, linked, or cited is functionally invisible , regardless of its historical importance.
The thesis that drives MIDF‘s work: unindexed knowledge becomes invisible knowledge.
What success looks like
The output of this work, if it succeeds, is not a website. It is a national knowledge layer.
The platform aims to unlock comprehensive ecosystem capabilities, ensuring that:

Image 9: One Knowledge Layer, Five User Groups. Source: MIDF Update II.
This is the shift MIDF is working toward: from scattered preservation to usable, interconnected infrastructure.
An invitation
The next phase is execution. MIDF will continue expanding the repository, strengthening metadata, onboarding manuscripts and inscriptions, supporting vulnerable collections, developing epigraphy workflows, and building toward full-text search, transliteration, OCR, and contributor systems , in stages, and with the discipline that the work requires. Poor scans, weak metadata, broken attribution, and non-standard formats create technical debt that compounds. The commitment is to build at a pace that is both credible and consequential.
MIDF invites support from four groups. Donors and CSR foundations can fund platform infrastructure, grants, storage, scanning, and metadata work. Scholars and universities can contribute to classification, transcription, review, and publication. Custodians and repositories can bring vulnerable collections into a durable digital pipeline, on their own terms. And technologists and volunteers can support OCR, search, transliteration, data cleaning, and contributor tooling.
India’s manuscripts and inscriptions have survived this long because generations of institutions, families, temples, mutts, scholars, and custodians protected them. The task now is different in kind: to make them discoverable, searchable, citable, and usable , for the scholarship of the next century.
That is the work MIDF has begun. This is an invitation to help scale it.
