LF AI & Data Foundation Launches DocLang Specification Working Group to Advance an Open Standard for AI-Native Documents

New specification, supported by leading LF AI & Data member organisations IBM and Red Hat, as well as others including ABBYY, complements the Docling open source project 

 LF AI & Data Foundation, the premier organisation supporting open source innovation in AI and data under the Linux Foundation, today announced the formation of the DocLang Specification Working Group. This working group supports a new collaborative standards development initiative to develop DocLang, an open, universal, AI-native document format designed to improve how enterprises prepare, exchange, and govern document data for AI systems. 

Founded by LF AI & Data premier members IBM, NVIDIA, and Red Hat, as well as contributors ABBYY and HumanSignal, the DocLang Working Group will operate under Joint Development Foundation’s vendor-neutral, open governance model to develop and maintain a specification that supports more reliable, interoperable document processing across AI and agentic workflows.

“Documents remain one of the most important sources of enterprise knowledge, but most were never designed for AI-driven workflows,” said Mark Collier, general manager of AI & Infrastructure at the Linux Foundation and Executive Director of LF AI & Data. “With the launch of the DocLang Working Group, we are bringing the open source community together to develop a vendor-neutral, interoperable standard that helps organisations prepare document data for AI more reliably, transparently, and at scale. Combined with projects like Docling, this effort can help create a more open foundation for document understanding across the AI ecosystem.”

“NVIDIA looks forward to working with the Linux Foundation and the broader DocLang ecosystem to accelerate the adoption of this AI-native document format across industries,” said Kari Briski, Vice President, Generative AI, NVIDIA.

Enterprises today work across a fragmented landscape of document formats, including PDFs, JPEGs, and other file types built primarily for human consumption rather than AI interpretation. As organisations increasingly rely on generative AI and agentic systems, this disconnect can introduce complexity, raise costs, and reduce reliability when extracting meaning from business documents.

“DocLang is designed to solve one of the foundational problems in enterprise AI: documents were built for humans, not machines,” said Maxime Vermeir, Vice President, AI Strategy at ABBYY. “By introducing a minimal, standardised, and AI-native representation of document structure, layout, meaning and governance, DocLang creates a far more deterministic foundation for modern AI systems. The results in an AI native context layer at scale.”

DocLang is designed to support:

  • Preservation of both semantic meaning and geometric layout in a single AI-native format
  • Representation of structural elements such as headings, paragraphs, and tables alongside their position on the page
  • Embedded governance controls to help downstream systems enforce policies related to privacy, extraction scope, and model training permissions
  • Optimisation for modern AI tokenization and modelling approaches to support more efficient and reliable document understanding

DocLang and Docling The new working group builds on the momentum of Docling, the open source document processing toolkit hosted by LF AI & Data. Originally developed by the AI for Knowledge team at IBM Research Zurich, and released as open source in 2024, Docling has become a widely adopted project for converting documents into structured, AI-ready representations.

Docling serves as the processing and conversion layer, ingesting a range of document formats (including .pdf, .docx, .pptx, .xlsx, HTML, and images) and transforming them into structured outputs using advanced models for layout analysis and table understanding. Its internal representation, DoclingDocument, captures text, tables, figures, reading order, and layout in a richly structured format.

DocLang complements that foundation by defining an open, interoperable standard for expressing and exchanging that structured output across systems. Together, Docling and DocLang create a more complete open source document AI stack under LF AI & Data, spanning document ingestion, parsing, standardized representation, and downstream consumption by language models and agentic AI systems.