---
title: "The indexing layer: making company data usable by AI"
url: https://automaark.com/insights/the-indexing-layer-making-company-data-usable-by-ai
canonical: https://automaark.com/insights/the-indexing-layer-making-company-data-usable-by-ai
description: "Most companies do not have a data problem. They have a findability problem. Before an agent, a search box or an analyst can use what you already have, something has to sit between raw storage and the question. That something is an indexing layer, and it is an architecture decision, not a plugin."
organization: Automaark
language: en
llms: https://automaark.com/llms.txt
author: Olamide Dada
author_url: https://automaark.com/authors/olamide-dada
date_published: 2026-09-29
date_modified: 2026-09-29
tags: Data platforms, AI systems, Data indexing, Retrieval
related_services: https://automaark.com/services/data-platforms, https://automaark.com/services/ai-agentic-systems
---

# The indexing layer: making company data usable by AI

Most companies do not have a data problem. They have a findability problem. Before an agent, a search box or an analyst can use what you already have, something has to sit between raw storage and the question. That something is an indexing layer, and it is an architecture decision, not a plugin.

## The short version

- Raw storage answers 'where is the record'. AI, search and analytics need 'what do we know about X', which is a different question with a different data shape.
- An indexing layer is the component that turns records into things you can ask about: entities resolved across systems, text made searchable, relationships made explicit, freshness tracked.
- Vector search is one index, not the layer. Structured filters, keyword search, entity resolution and access control are the parts people skip and then regret.
- Design the layer once, independent of any model or vendor. Models change monthly; the shape of your company's knowledge does not.

Every company we talk to about AI has the same first sentence: "We have all this data." It is true. It is also the reason the first AI project stalls.

Having data is not the same as being able to ask it a question. A CRM has records. A ticketing system has records. A shared drive has ten thousand documents and a naming convention that lasted three weeks. None of these can answer "what do we know about this customer's history with us?" without a person opening four tabs and remembering things.

An AI system cannot open four tabs. It needs something to ask. Building that something is what we mean by the indexing layer, and it is the part of most AI projects that determines whether the demo becomes a product.

## Storage answers the wrong question

Operational systems are built to store and retrieve *records*: this order, that ticket, the contract with this ID. They are very good at it. The question they answer is "where is the thing?"

The questions people, search and agents actually ask are different in kind:

- "What do we know about Acme?" (an entity, spread across six systems, under four spellings)
- "Which contracts mention a termination clause shorter than 30 days?" (meaning, not keywords)
- "Who handled the last three escalations from this account, and what did they promise?" (relationships and history)
- "Show me everything I am allowed to see about this project." (permissions, which live nowhere central)

None of these has a home in a transactional schema. You can bolt a search box onto one system, but the moment the answer spans two, you are writing glue, and glue is where AI projects go to die.

## What the layer is

An indexing layer is a component whose only job is to turn records into *things you can ask about*. Concretely it does five things, and skipping any of them produces a specific, predictable failure later.

### 1. Entity resolution

"Acme Ltd", "ACME", "Acme Limited (UK)" and the domain `acme.co.uk` are one customer. The layer decides that once, assigns a stable identifier, and links every record from every source to it. Without this, every query is a fuzzy match, and every answer is partial.

### 2. Text made searchable, in more than one way

Documents, emails, tickets and notes need to be findable by exact terms (an invoice number), by meaning (a complaint about delivery, however it is phrased) and by structure (only PDFs, only last quarter). That is at least three indexes over the same content: keyword, vector and structured. Vector search alone is the most common mistake we see; it is excellent at "things like this" and useless at "invoice 4471".

### 3. Relationships made explicit

Customer → contracts → invoices → disputes → the people involved. Transactional systems hold these as foreign keys inside one database. Across databases the links do not exist until you build them. This is the graph, and it is what lets an agent answer a multi-step question without a human stitching the steps.

### 4. Freshness and lineage

Where did this fact come from, when was it last true, and has the source changed since? An index that cannot answer this will confidently serve stale data, and stale data from an AI system is worse than no answer, because it arrives with a fluent explanation.

### 5. Access control at the index

If permissions live only in the source systems, the index becomes a way around them. Every indexed item carries who may see it, and every query is filtered before retrieval, not after. This is the requirement that most turns a prototype into a rebuild when it is discovered late.

## Why it is architecture, not a plugin

The vendor pitch is that you connect your sources to a product and the layer appears. For a single source and a simple question, that is roughly true. It stops being true at the second source, because entity resolution, relationships and permissions are *about your business*, and no product knows your business.

It is also why we design the layer independently of any model or vendor. The models are changing monthly. The stores are changing yearly. The shape of what a company knows, its entities, its relationships and who may see what, changes slowly, and it is the expensive part to get wrong. Design that once, as a contract, and let the models and stores behind it be swapped.

The practical consequence: an indexing layer is a data-platform engagement first and an AI engagement second. The AI is the consumer. If the layer is right, the first agent takes weeks to add and the fifth takes days. If the layer is missing, every agent rebuilds a private, inconsistent version of it, and the company ends up with six different answers to "what do we know about Acme?"

## What "done" looks like

We use a short checklist to decide whether a layer is real or still a prototype:

1. One stable identifier per entity, resolved across every source that mentions it.
2. A query for "everything about X" that returns in under a second and respects the asker's permissions.
3. Keyword, meaning and structured filters combinable in a single request.
4. Every result carries its source, timestamp and a link back to the system of record.
5. A new source can be added without changing any consumer.

Hit those five and the interesting work starts: agents that act, search that people trust, analytics that do not need a data team to run a query. Miss them and you have a chatbot that is impressive for exactly one demo.

## The honest position

The indexing layer is unglamorous. It does not appear in the pitch deck. It is also the part of the system that most decides whether AI does real work for a company or stays a novelty, which is why we treat it as one of our seven departments rather than a feature of another one.

If you have "all this data" and the first AI project stalled, this is almost certainly why.

## Questions this raises

### Is an indexing layer the same as RAG?

RAG (retrieval-augmented generation) is one consumer of an indexing layer: it retrieves passages and hands them to a model. The layer itself is broader: it holds resolved entities, structured fields, relationships and permissions, and serves search, analytics and agents as well as generation. Building 'just RAG' usually means rebuilding the layer three times as each new use case arrives.

### Do we need a vector database?

Sometimes. A vector index is useful when the question is about meaning rather than exact terms. But most business queries are hybrid: 'invoices over $10k for this customer that mention a dispute' needs structured filters, keyword matching and access control before similarity search adds anything. Pick the store after the layer is designed, not before.

### How long does it take to build?

It depends on how many source systems there are and how clean their identifiers are. The scoping phase produces a written answer for your case, with milestones, before any build work starts.

## Sources

- [Designing Data-Intensive Applications, chapter 3: Storage and Retrieval (Kleppmann)](https://dataintensive.net/)
- [Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020)](https://arxiv.org/abs/2005.11401)

---
Written by Olamide Dada, Founder & CEO, Automaark. More: https://automaark.com/insights
