Thousands of pipelines a month
Thousands of pipelines a month
1 use case live
1 use case live
Zero data retention
Zero data retention
CUSTOMER STORY / Data automation
CUSTOMER STORY / Data automation

Helping Maia build thousands of data pipelines

“When Maia reaches for search_documentation, speed and accuracy are essential. Kapa always responds within a few seconds and the responses contain exactly the right context for Maia. Crucially, Kapa’s pruning feature also ensures that Maia doesn’t receive any irrelevant context.”

Sam Perrin

Principal AI Engineer, Matillion

Challenge

Maia builds and maintains thousands of data pipelines for Matillion users every month, and many of its tasks depend on knowledge of Matillion’s own products: which component to use, what a configuration expects, how a connector behaves. Even almost-accurate product knowledge steers Maia down a wrong path and ends in broken pipelines and frustrated customers.

Solution

Give Maia a search_documentation tool powered by Kapa’s retrieval engine. Each call is interpreted and decomposed, matched against a live index of Matillion’s documentation, pruned of stale or irrelevant content, and returned as a handful of source-linked passages—giving Maia current answers it can trust before changing a pipeline. For Matillion, it is a single MCP-formatted API call with zero data retention.

Results

  • Thousands of data pipelines built and maintained by Maia every month

  • Responses within a few seconds on every search_documentation call

  • Exactly the right context, with irrelevant passages pruned before they reach Maia

  • Zero data retention across the integration

  • Zero engineers pulled off Maia to build retrieval in-house

THE SITUATION

An agentic data engineer that needs product facts mid-task

Maia is an AI Data Automation platform that uses autonomous agents to build and maintain data products and pipelines. It builds and maintains thousands of data pipelines on behalf of Matillion users every month, and many of the tasks it receives require knowledge of Matillion’s own products:

  • Which component to use

  • What a configuration expects

  • How a connector behaves

Raspberry Pi is a full-stack engineering company. It designs its own silicon, boards, and operating system, and it serves three very different audiences:
Industrial and embedded customers
Enthusiasts and educators
Semiconductor buyers
Each group comes to the documentation site with different use-cases and questions.
The core documentation at raspberrypi.com/documentation is hosted on a website, but a large share of the most valuable technical content is not. The Product Information Portal (PIP) holds the product briefs, datasheets, and whitepapers that industrial customers rely on, almost all of those are PDFs. Raspberry Pis Official Magazine publishes issues that run to hundreds of pages each, also as PDFs. Some tooling, such as rpi-image-gen, documents itself inside its GitHub repository in code.
Keyword search on the docs site only surfaced the core documentation, and as Gordon put it in the announcement post: “with keyword search you have to guess the author's vocabulary.”
Gordon had already prototyped a retrieval-augmented generation (RAG) system of his own, so he understood what the technology could do and what it would take to run one in production. The question was whether to keep building or to buy, and if buying, from which vendor.

Even almost-accurate product knowledge steers Maia down a wrong path and results in broken data pipelines and frustrated customers.

So Maia needs a reliable way to look things up mid-task, at the volume thousands of users generate.

Raspberry Pi is a full-stack engineering company. It designs its own silicon, boards, and operating system, and it serves three very different audiences:
Industrial and embedded customers
Enthusiasts and educators
Semiconductor buyers
Each group comes to the documentation site with different use-cases and questions.
The core documentation at raspberrypi.com/documentation is hosted on a website, but a large share of the most valuable technical content is not. The Product Information Portal (PIP) holds the product briefs, datasheets, and whitepapers that industrial customers rely on, almost all of those are PDFs. Raspberry Pis Official Magazine publishes issues that run to hundreds of pages each, also as PDFs. Some tooling, such as rpi-image-gen, documents itself inside its GitHub repository in code.
Keyword search on the docs site only surfaced the core documentation, and as Gordon put it in the announcement post: “with keyword search you have to guess the author's vocabulary.”
Gordon had already prototyped a retrieval-augmented generation (RAG) system of his own, so he understood what the technology could do and what it would take to run one in production. The question was whether to keep building or to buy, and if buying, from which vendor.

search_documentation: one tool call, powered by Kapa

Ask Maia to move Salesforce data into Snowflake and it needs facts fast: how that connector authenticates, which components handle incremental loads, what each configuration expects. For those moments, one of the 10+ tools in Maia’s toolkit is search_documentation, powered by Kapa. Each call runs through Kapa’s retrieval engine: the query is interpreted and decomposed, matched against a live index of Matillion’s documentation, component references, guides, and more, pruned of stale or irrelevant content, and returned as a handful of source-linked passages Maia can work from mid-task. From Matillion’s side, the integration is just an MCP-formatted API call, and it respects Matillion’s zero data retention requirements.

Play video

Why not build it themselves?

The short answer: opportunity cost. Matillion’s AI team exists to make Maia a better data engineer. Every week spent on chunking strategies, ranking experiments, index maintenance, and pruning is a week taken from the core goal. Meanwhile, state-of-the-art retrieval is Kapa’s core goal.

“We could build agentic search ourselves. But every engineer we’d put on it is an engineer taken off Maia. Kapa’s cost and performance made it an easy decision.” Sam Perrin, Principal AI Engineer, Matillion

THE RESULT

Speed and accuracy at the volume thousands of users generate

When Maia reaches for search_documentation, speed and accuracy are essential. Kapa responds within a few seconds, and the responses contain exactly the right context for Maia to keep building without a wrong turn.

Three things matter to Matillion on every call:

Fast. Every search_documentation call returns within a few seconds, so a documentation lookup never stalls a pipeline build.

Precise. Responses contain exactly the right context, and Kapa’s pruning strips stale or irrelevant passages before they reach Maia.

Private. The integration is an MCP-formatted API call that respects Matillion’s zero data retention requirements.

For Matillion’s AI team, the result is focus: state-of-the-art retrieval without weeks spent on chunking strategies, ranking experiments, index maintenance, or pruning, so every engineer stays on making Maia a better data engineer.

Everything you need to give your agents accurate retrieval, ready in minutes.

Get started for free. Connect your sources and get an API key or hosted MCP server to give your agents the context they need to do reliable work at scale.