Publication
Knowledge Factory for the AI-Native Organization: How the GAP Framework Turns the Chaos of Raw Sources into a Trusted Canon
This article was written specifically for DOU.
A logical sequel to the article "The Third-Generation Ceiling: Why Organizations Need a Dual-Circuit Knowledge Architecture in the AI Era". In the first part, we established the division of the information space into Operational and Institutional circuits. Here, we analyze the production mechanism itself: exactly how knowledge is born, verified, and anchored in an organization's canon. We examine the GAP (Grounded Agentic Pipeline) framework, compare Runtime RAG with Build-time GAP, break down the 5 quality control stages, explore multi-format delivery, and share practical insights from real-world public projects.
1. From Information Anatomy — to a Knowledge Production Line
In the previous article, I showed why modern companies have hit the "ceiling of 3.0 Organizations." Unconstrained communication across Slack or Teams and the chaos of shared cloud drives created chronic entropy: documentation becomes obsolete at lightning speed, decision context vanishes during employee turnover, and instead of reliable instructions, dozens of unsynchronized drafts pile up.
The solution was dividing information into two spaces: the Operational Circuit (where people communicate freely, experiment, and draft hypotheses) is strictly separated from the Institutional Circuit (where the company's verified digital canon lives) through a controlled Promotion Gate.
[ OPERATIONAL CIRCUIT ]
(Drafts, hypotheses, discussions)
│
▼
┌───────────────────┐
│ GAP │ ◄── HOW EXACTLY IS KNOWLEDGE
│ AND GATE │ CREATED & VALIDATED HERE?
└─────────┬─────────┘
│
▼
[ INSTITUTIONAL CIRCUIT ]
(Trusted canon: Git, Markdown, SSOT)
The key practical question that arises immediately after is:
How exactly should knowledge be produced within this system?
Who takes hundreds of disparate raw sources — regulations, meeting transcripts, industry standards, expert interviews — and transforms them into flawless playbooks, training courses, encyclopedias, and strategies that guide the organization for years?
Three Failures of Uncontrolled Generation
It might seem that the easiest shortcut is to assign a language model (or a cluster of agents), give it access to storage, and prompt it to write policies. But without a rigorous engineering harness, such generation consistently leads to three critical failures:
- Hallucinations & Fact Drift: The model easily fabricates details whenever the source texts have gaps. Worse, between different working sessions, the agent contradicts itself, producing differing versions of the "truth" every time.
- Loss of Systemic Consistency: Terminology, document structure, and tone of voice begin to drift. What the model formulated on Tuesday clashes in terms with what another agent or human colleague authored on Thursday.
- Lack of Auditability and Accountability: When a new document appears in storage or an existing one changes, it is virtually impossible to trace which exact source triggered a given phrasing or whether a domain specialist ever vetted it.
Without a structured technological pipeline, bringing in AI only accelerates the pollution of documentation with plausible-sounding yet unverified texts. To overcome this barrier, an engineering discipline is required — the GAP (Grounded Agentic Pipeline) framework.
2. Paradigm Shift: Runtime RAG vs. Build-Time GAP
The prevailing approach today for connecting AI to corporate knowledge is the RAG (Retrieval-Augmented Generation) pattern: documents are chunked, stored in a vector database, and for every user query, the system searches for similar snippets, passing them to the model at prompt generation time.
For quick one-off answers, this works. But for establishing an organization's foundational knowledge, classic RAG has a fatal flaw: it defers all processing, error filtering, and conflict resolution to Runtime, dumping the cognitive risk directly onto the end user.
| Comparison Criterion | Runtime RAG (Search at query time) | Build-time GAP (Knowledge manufacturing line) |
|---|---|---|
| When processing occurs | The moment the user hits "Enter" and waits for a response. | Upfront, during the knowledge structuring, validation, and release phase (Build-time). |
| Disinformation risk | Constant: The user sees live synthesis; if contradictory chunks are retrieved, the error appears right before their eyes. | Minimized before release: Facts and wordings are cross-checked and reviewed before publication. |
| Contextual integrity | Fragmented: Vector search pulls isolated paragraphs, frequently stripping out macro structure and causal links. | Systemic: Knowledge is synthesized into coherent, comprehensive articles preserving context. |
| Speed and cost | Retrieval and inference latency; continuous token burning on every single click. | Instant access to structured text; zero token costs for standard reading of the canon. |
| Auditability & Traceability | Difficult to recreate why a model retrieved specific chunks six months ago. | Full Git history: clear view of the exact raw source, author, and timestamp of every commit. |
[ TRADITIONAL RAG: Risk on the end user ]
Raw sources ──► Vectorization ──► [User query] ──► Chunk search ──► On-the-fly generation ──► End-user risk!
[ GAP FRAMEWORK: Quality guaranteed on the pipeline ]
Raw sources ──► [GAP Pipeline: Agent + Constitution + Review] ──► Verified Canon ──► Instant trusted access
Crucial nuance: GAP does not wage war against RAG. On the contrary: the structured and interlinked output of GAP often makes standalone vector RAG unnecessary — humans and agents alike can find precise answers directly in compiled files. Where vector search is genuinely needed (e.g., across hundreds of thousands of pages), GAP provides it not with noisy chaff, but with clean, verified, highly structured context.
3. The Architectural Core of GAP: 5 Quality Assurance Stages
The concept of GAP (Grounded Agentic Pipeline) solves a dual challenge: it bridges the gap between raw unstructured texts and polished documentation, while keeping the agents' work strictly grounded in an empirical, factual foundation.
The GAP pipeline consists of 5 sequential stages:
┌────────────────────────────────────────────────────┐
│ 1. SOURCE OF TRUTH │
│ • Immutable primary materials (raw/) │
│ • Grounding facts, zero hallucinations │
└─────────────────────────┬──────────────────────────┘
│
▼
┌────────────────────────────────────────────────────┐
│ 2. CONSTITUTION │
│ • Rules of engagement & standards (AGENTS.md) │
│ • Structure conventions, templates, and schemas │
└─────────────────────────┬──────────────────────────┘
│
▼
┌────────────────────────────────────────────────────┐
│ 3. AGENT & SKILLS │
│ • Orchestrator + isolated subagents │
│ • Specialized skills (.agents/skills/) │
└─────────────────────────┬──────────────────────────┘
│
▼
┌────────────────────────────────────────────────────┐
│ 4. REVIEW GATE │
│ • Cross-checking with sources & link linting │
│ • Audit log (log.md) + Human-in-the-Loop │
└─────────────────────────┬──────────────────────────┘
│
▼
┌────────────────────────────────────────────────────┐
│ 5. DELIVERY │
│ • Verified canon (Markdown + YAML) │
│ • Obsidian, web portals, PDF reports, API │
└────────────────────────────────────────────────────┘
Stage 1. Source of Truth (raw/)
Everything starts with primary source materials: legislation, technical standards, meeting transcripts, policy documents, or interviews. All of them reside in a dedicated raw/ or references/ directory.
- Immutable Raw Rule: Once source material is placed in a date-stamped folder, any modification to it is strictly prohibited.
- Zero Hallucination Mandate: The agent is explicitly forbidden from guessing based on the model's internal parametric memory or performing unguided web browsing. The primary source is the sole arbiter of facts.
Stage 2. Project Constitution (AGENTS.md)
Agents cannot operate on verbal agreements or informal expectations. Alongside the content sits a version-controlled rules document — the Constitution (AGENTS.md or CLAUDE.md).
- This serves as the permanent engineering law of the repository, committed to version control.
- The Constitution specifies folder structure, metadata formatting (YAML frontmatter), mandatory document templates, file naming rules (Latin characters in
kebab-case), relative link conventions, and explicit constraints (e.g., prohibitions against direct commits tomainor altering historical audit entries).
Stage 3. Executing Agent and Modular Skills (Agent + Skills)
Rather than cramming all instructions into a single sprawling prompt, operations are partitioned into reusable instructions — Skills in the .agents/skills/ directory:
- A skill for source ingestion and initial parsing (
ingest); - A skill for synthesizing canonical articles or policies (
create); - A skill for verifying relational links and schema metadata (
linter).
Subagent Pattern and Context Isolation:
Stuffing 50 large reports into a single context window inevitably degrades model performance: it loses instructions, mixes up dates, and starts hallucinating. GAP resolves this via delegation to isolated subagents:
- The lead orchestrator breaks down the objective into discrete subtasks ("Analyze Section 4 of the standard and extract safety requirements").
- A subagent launches with a pristine context window, takes only the relevant slice of the source text, executes the skill, and returns a structured output.
┌────────────────────────────────────────────────────────┐
│ MAIN ORCHESTRATOR (Lead Agent) │
│ • Guided by the Constitution (AGENTS.md) │
│ • Keeps its own context clean of raw dumps │
└──────────────┬──────────────────────────┬──────────────┘
│ │
[Task Delegation 1] [Task Delegation 2]
▼ ▼
┌──────────────────────────┐ ┌──────────────────────────┐
│ SUBAGENT A (Ingest) │ │ SUBAGENT B (Synthesis) │
│ • Clean context window │ │ • Clean context window │
│ • Focused on Source #1 │ │ • Focused on Source #2 │
└──────────────────────────┘ └──────────────────────────┘
Stage 4. Review Gate & Quality Control
An agent's output is never assumed to be accepted by default.
- Grounding Check: Every substantive assertion must include a backlink or provenance tag referencing the input file in
raw/. - Automated Linting: Scripts verify the validity of relative links, ensure mandatory frontmatter fields exist, and detect orphaned pages.
- Audit Log (
log.md): Every action is appended to the bottom of the log file in append-only mode. Modifying or deleting past entries is forbidden. - Human-in-the-Loop: At designated checkpoints, a human subject-matter expert reviews a standard
git diffand approves the pull request.
Stage 5. Delivery & Multi-Format Consumption
Resulting knowledge is not locked inside proprietary software silos. It is stored in a clean, vendor-neutral format — pure Markdown with YAML metadata, accessible by any tool in the company's ecosystem.
4. Multi-Format Consumption: 5 Scenarios for Using the Canon
A cornerstone principle of this architecture is the Separation of Content and Presentation.
A single, verified body of knowledge concurrently powers multiple organizational workflows:
┌──► 1. Local Knowledge Base (Obsidian Vault)
│
├──► 2. Corporate Web Portal (MkDocs, Docusaurus)
VERIFIED CANON │
OF GAP KNOWLEDGE ──────┼──► 3. Export to Artifacts (PDF, policies, slides)
(Markdown + frontmatter) │
├──► 4. Input for Next Pipeline (Chaining)
│
└──► 5. Direct Agent Search (grep, metadata)
- Interactive Knowledge Base (Obsidian Vault):
The project repository can simply be opened as an Obsidian vault. Thanks to standardized internal links, the team immediately gets a live Knowledge Graph, backlinks, and instant search without provisioning databases or server infrastructure. - Deterministic Publishing (Build & Publish):
The repository connects to static site generators (such as MkDocs, Docusaurus, Astro, or Quartz). Merges to the release branch trigger standard CI/CD pipelines that compile markdown into a fast, navigable documentation site. The golden rule: only source Markdown changes; generated HTML is never modified directly. - Export to Regulatory and Operational Artifacts:
Automated transformation scripts (via Pandoc or WeasyPrint) compile canon files into formalized artifacts. As needed, the system generates official directives, PDF policy packets, or executive slide decks. - Pipeline Chaining:
The verified output of one pipeline serves as an immutable source for a subsequent, narrower workflow:- Pipeline 1: Processing an industry standard ──► Generating the organization's internal playbook.
- Pipeline 2: Internal playbook as source ──► Synthesizing an employee training curriculum.
- Pipeline 3: Curriculum as source ──► Automatically generating certification exam questions.
- Direct Agent Search in the Knowledge Base:
Because knowledge is stored as plain text with clean semantic headings and YAML tags, AI assistants rarely require vector databases. Agents locate precise statements using standard keyword and metadata search with zero risk of context contamination.
5. Engineering Foundation: Docs-as-Code and the Power of Git
Adopting Markdown + Git instead of yet another no-code database is a pragmatic choice. Software engineering tooling has been honed for decades specifically to manage changes in complex textual systems:
- Granular Line-by-Line Review (
git diff): A reviewer never has to re-read an entire document. They inspect an exact diff: which words the agent added or removed, which links were updated, and what new claims emerged. - Isolated Experiments via Branches: Substantial updates to a policy or strategy are developed in isolated branches. The main canon remains untouched until all changes pass validation.
- Knowledge Review Culture (Pull Requests): Merging any new article or modification into the canon follows standard code review rituals, adapted for knowledge governance.
6. Reference Repository Architecture for GAP
To preserve order, the repository adheres to a strict directory hierarchy:
.
├── AGENTS.md (and CLAUDE.md) # CONSTITUTION: architecture, constraints, agent rules
├── .agents/skills/ # SKILLS CATALOG: reusable scripts & procedure instructions
│ ├── ingest/SKILL.md # Source ingestion and initial analysis procedure
│ ├── linter/SKILL.md # Link integrity and metadata validation procedure
│ └── query/SKILL.md # Synthetic analytical report generation procedure
│
├── inbox/ # Staging buffer for raw files before committing
│ └── assets/ # Staged media files (images, diagrams)
│
├── raw/ # 1. IMMUTABLE RAW SOURCES (Source of Truth, Read-Only)
│ ├── YYYY-MM-DD/ # Archive of sources organized by ingestion date
│ └── assets/ # Archived attachments and media from sources
│
├── templates/ # TEMPLATES: required scaffolds for concepts, entities, logs
│ ├── article.md
│ ├── index.md
│ └── log.md
│
└── wiki/ (or policy/, course/) # 2. VERIFIED KNOWLEDGE (Pipeline Output / Canon)
├── concepts/ # Core concepts, theoretical models, technical standards
├── entities/ # Key figures, organizations, systems, tools
├── archives/ # Saved comprehensive analytical query responses
├── index.md # Master navigation index of all knowledge units
└── log.md # Append-only audit log (who, when, and what changed)
This layout ensures complete audit transparency: any assertion in wiki/concepts/ can be traced in seconds via a corresponding entry in log.md directly back to a specific paragraph in raw/YYYY-MM-DD/.
7. Evidence Base: Public Projects Across Diverse Domains
The GAP framework emerged from the synthesis of practical hands-on experience building structured knowledge repositories.
The original inspiration came from the LLM Wiki concept proposed by OpenAI co-founder Andrej Karpathy. Karpathy outlined using large language models to maintain a personal accumulating wiki. I took that concept and systematized it for rigorous organizational production: adding a Project Constitution, subagent context isolation, Review Gates, audit logging, and multi-format delivery.
This approach was detailed in the article "From RAG to LLM Wiki", which won Best Technical Article of July on DOU. The reference starter template is published in the open-source repository llm-wiki.
Here are several public projects built upon these principles (others operate within closed, proprietary environments):
| Project & Link | Domain & Input Sources (raw/) |
Pipeline Output (wiki/ / Delivery) |
|---|---|---|
| 1. "Crimea is Ukraine" Encyclopedia: wiki.crimea-is-ukraine.org |
Historical & socio-political knowledge base. Sources: 200+ disparate articles from the crimea-is-ukraine.org portal. |
Encyclopedia with 400+ articles: Structured static portal with automatically interlinked concepts, historical events, personalities, and full-text search. |
| 2. "Scientific Image of the World" Portal: scientific-image.bogdanovych.org |
Academic physics & scientific worldview. Sources: 50+ transcripts of lecture series by KNU Associate Professor M. Vysotskyi. |
Scientific educational portal (50+ articles): Structured lecture outlines, unified glossary of terms, cross-references to laws of physics, and scientist biographies. |
| 3. Personal Health Analytics Base: mental-health |
Personal health monitoring & analytics. Sources: Daily structured logs, questionnaires, subjective metrics. |
Local Markdown vault + automated reporting: Weekly and monthly condition dynamics, correlation analysis without leaking private data to third-party clouds. |
Across diverse domains — from historical encyclopedias to personal analytics — a shared truth emerges: knowledge reliability rests on an unbroken chain: Source ──► Constitution ──► Agent ──► Review ──► Delivery.
8. Where GAP Is Not Suitable: Boundaries of Application
Every tool has its optimal domain, and GAP is no exception. Attempting to force it everywhere indiscriminately only introduces unwarranted overhead.
| Where GAP Fits | Where GAP Does Not Fit |
|---|---|
| • Regulatory acts, policies, operating standards | • Real-time data (stock prices, telemetry, live metrics) |
| • Training courses, textbooks, certification tests | • Ad-hoc one-off lookups in chat format |
| • Corporate knowledge repositories & encyclopedias | • Projects lacking trusted primary sources |
| • Long-term architectural and strategic reports | • Teams without review and versioning habits |
- Human Review Bottleneck:
GAP does not eliminate the human from the loop; it shifts human attention from query response time to the knowledge compilation stage. On high-stakes documents, an expert must still inspect proposed changes. - Static Snapshots:
The approach is designed for periodic compilation releases. It is not built for streaming, ephemeral data feeds — like live server monitoring or exchange ticks. - Tooling Skill Barrier:
Docs-as-Code requires basic familiarity with Git, branching concepts, and Markdown formatting. Without friendly abstraction layers for non-technical teams, onboarding friction can be significant. - GIGO Principle (Garbage In — Garbage Out):
If primary sources are incomplete, fabricated, or fundamentally contradictory, the pipeline cannot transform them into a reliable canon. - Need for Constitution Maintenance:
If business guidelines evolve while theAGENTS.mdfile remains untouched, agents will continue generating documentation against outdated specifications.
9. Launch Checklist: 6 Practical Steps for the Team
A step-by-step roadmap to spinning up your own knowledge factory:
- Step 1. Anchor primary sources (
raw/): Store all authoritative source documents in a dedicated, immutable folder and strictly restrict agents to using only these texts. - Step 2. Author the Project Constitution (
AGENTS.md): Define repository layout, frontmatter requirements, naming conventions, and prohibited actions. - Step 3. Carve out initial skills (
.agents/skills/): Start with two straightforward operations —ingestto import and parse new materials, andlintto validate links and structure. - Step 4. Establish Review Gate & Audit Log (
log.md): Mandate reviewing changes viagit diffbefore merging, and record each pipeline action in an append-only log. - Step 5. Define delivery targets: Choose how team members will consume the verified material: locally via Obsidian, via a static MkDocs web portal, or through automated PDF export.
- Step 6. Execute a pilot sprint: Run one focused document or policy section through the complete pipeline, test the review dynamics, and calibrate rules before scaling across the organization.
10. Conclusion: Institutional Memory Instead of Stochastic Answers
Combining both architectural concepts yields a pragmatic operating framework for organizations in the AI era:
- The Dual-Circuit Model preserves information hygiene: drafts, brainstorms, and fast-moving dialogues remain in the Operational Circuit, while validated decisions migrate to the Institutional Circuit.
- The GAP Framework acts as the manufacturing engine: taking raw texts and, through subagents, the Constitution, and verification gates, transforming them into resilient digital assets.
Instead of hoping for an accidental stroke of prompt-engineering luck in a chatbot each time, building a repeatable engineering discipline is vastly more powerful. When every assertion is grounded in a primary source and every revision is preserved in Git, an organization's knowledge base becomes a managed institutional asset rather than an unverified heap of AI responses.