K-Lake KDBL Context Lake

Give the AI you already use
the documents you already have.

Works with Microsoft Copilot Claude ChatGPT your own agents

K-Lake is not another chatbot. It connects the assistants you already pay for to the files in your shares, SharePoint and cloud storage. Nothing is copied or migrated, nobody gains access they do not already have, and every answer shows the document it came from.

No migration. Files stay where they are. Runs in your own environment Up to 7x fewer AI tokens per answer Every answer shows its source
Microsoft CopilotClaudeChatGPT via K-Lake priya@northwind.com
Which supplier agreements renew in the next 90 days, and who signed them for us?
search_content 3 sources · 1,204 candidate files trimmed to what Priya can open

Three agreements in your readable documents renew before 7 December:

  • Meridian Logistics GmbH, master services agreement, renews 30 October. Signed by James Whitfield. msa-meridian-2024.pdf · p.14
  • Alder & Finch LLP, advisory retainer, renews 18 November. Signed by Priya Raman. retainer-alder-finch.docx
  • Cormorant Freight BV, data processing addendum, renews 2 December. Signature block names James Whitfield. dpa-cormorant-signed.pdf · p.9
Every fact cites its file Query recorded in the audit trail Nothing left your cluster

Illustrative exchange: the same answer from Copilot, Claude or ChatGPT, through K-Lake. Names and files are fictional.

Works with the assistants you already use
Microsoft CopilotClaudeChatGPTCursorYour own agentsAny MCP client
Reads the files where they already live
SharePointOneDriveWindows file sharesNFSAmazon S3Azure Blob
Signs in with the identity you already run
Microsoft Entra IDOktaGoogleKeycloakAny OIDC provider

Not another chatbot. The link between the AI you use and the files you own.

K-Lake has no assistant of its own. It gives Copilot, Claude, ChatGPT and your own agents a safe, governed way to read the contracts, reports, spreadsheets, scans and recordings you already hold, without moving a single file, and makes every answer cheaper to produce.

Keep control. Nothing moves.

K-Lake runs in your own environment and reads your files where they already live. Nothing is copied into a vendor's platform, nothing is migrated, and nothing is sent to a model provider unless you choose it. Air-gapped sites are fully supported.

How it runs in your environment

Use the AI you already bought.

Microsoft Copilot, Claude, ChatGPT, Cursor or agents you build yourself all connect the same way, through the open standard they already speak. Switch assistants or run several side by side. No lock-in to one model vendor.

Connecting your assistant

Spend less on every answer.

Tokens are the meter on every AI bill. An assistant left to hunt through folders burns them opening file after file. K-Lake hands it the right passages first time. In our testing that used 2.5x fewer tokens than an assistant searching the files itself, and 7x fewer than a standard vector store.

Where the saving comes from

Answers you can trust and defend.

People and assistants only see what they could already open, using the permissions your files carry today. Every answer shows the file, page or minute each fact came from, and every question is recorded so you can show an auditor who asked what.

Security and audit

The same answer, for a fraction of the tokens.

Every question an assistant answers is metered in tokens, and tokens are what every AI bill is made of. How the assistant finds your documents decides how many it burns.

Assistant without K-Lake

Reads far more than it needs to find one paragraph.

  1. Searches file by file for likely words, or pulls fixed-size fragments from a vector store
  2. Reads whole documents, or goes back for the context a fragment left out
  3. Discards the matches that were near the topic but not the answer
  4. Answers, and re-sends everything it read on every later turn

Every step is billed

7x fewer tokens, up to

In our testing, K-Lake used 2.5x fewer input tokens than an assistant searching the files itself with grep and command-line tools, and 7x fewer than a standard vector-chunk store.

Assistant with K-Lake

Asks once. Gets the right passages back.

  1. Asks K-Lake, which has already read every file
  2. Receives the relevant passages, each with its source
  3. Answers

One round trip

Lower monthly AI bills

Fewer tokens per question means the same usage costs less, whichever assistant or model you run.

Faster answers

One retrieval instead of a dozen file reads. People get the answer while they still remember the question.

Licences and limits go further

Rate limits and per-seat allowances stretch across more questions, so a rollout scales without a surprise on the invoice.

Measured by KDBL in August 2026 on 40 public company filings and three research questions. K-Lake usage was observed; the two alternatives were modelled on the same tasks. Savings vary with the assistant, the model and the shape of your files. Read the method, numbers and limits.

One index. Four ways to put it to work.

Crawl, extract and enrich once. Then search it, ask it, explore it, or act on it, from the console, the command line, the API, or an assistant.

Grounded AI over MCP

Put any assistant on top of your own content, safely.

K-Lake ships a Model Context Protocol server, so Claude, ChatGPT, Cursor, IDE assistants and your own agents can search and read your extracted content without a bespoke integration. It is a full OAuth 2.1 authorization server that federates sign-in to your identity provider.

  • Assistants connect as the individual user and inherit their per-file trimming
  • A live access flow shows who is asking, from which client, hitting which sources
  • Revoke a session or hide a source from AI with one click
See the MCP server
The K-Lake MCP console showing a live access flow from users, through AI platforms and tools, to the sources they reached

The MCP access flow in the K-Lake console: users, platforms, tools and sources, live.

Knowledge graph

Who and what is in our documents?

Search answers "which files mention this?". The knowledge graph answers the questions that come next. K-Lake reads the text it has extracted and records the people, organisations, locations and agreements it names, and the relationships the text asserts between them.

  • Ask which files put three people in the same room, in one query
  • Tell asserted relationships apart from mere co-occurrence
  • Merge duplicate spellings of a name, reversibly
Explore the graph
The Explore page in K-Lake showing an organisation at the centre of an orbit of related people, agreements, counterparties and locations

Orbit view around one organisation. Solid edges are asserted relations; dashed edges are co-occurrence.

Discover

What data do we actually have?

A data-estate overview built from what K-Lake has already crawled. See what exists, where it lives, how old it is, how much is searchable and how exposed it is, and click straight through from a percentage to the files behind it.

  • Composition by file type, content type, language, size, age, owner, storage tier and source
  • Exposure signals: world-readable, stale, unowned and duplicated files
  • The same view for an AI assistant, so it can orient before it searches
See Discover
The Discover dashboard in K-Lake with estate totals, a searchability bar and composition breakdowns by file type, content type and language

Discover: estate totals, searchability and composition, with a governance view below.

Smart Actions

Act on every file as it is found.

After each crawl, every file flows through a pipeline you compose per source. Extract its content, redact identifiers and card numbers from what search and AI return, notify your own systems with a signed webhook, or retry hard documents on a heavier engine.

  • Change a rule and only that step re-runs, with no re-crawl
  • Add a step to an existing source and it backfills automatically
  • Everything derived is stored alongside the file, never written back into it
See Smart Actions
The Smart Actions pipeline on a source: content extraction followed by redaction, with controls to add, reorder and configure each step

A source's pipeline: extraction, then redaction, then whatever you add next.

A read-only layer over the systems that already own your files.

Nothing is moved, renamed or rewritten. K-Lake reads your sources, builds an index in your cluster, and serves it through four interchangeable interfaces.

Your sources

Object storesAmazon S3, Azure Blob, S3-compatible
File sharesSMB / CIFS, NFS, with their ACLs
Microsoft 365SharePoint libraries, OneDrive
Identity providerEntra ID, Okta, Google, Keycloak
read-only

K-Lake

Runs inside your Kubernetes cluster. No egress required.

CrawlInventory every file, its size, age, hash and native permissions
ExtractText, tables and OCR from documents; transcripts from audio and video
Enrich and actEntities and relations, redaction, webhooks
Index and trimHybrid search with per-tenant and per-file row-level security
trimmed and audited

Your people and agents

AI assistants over MCPClaude, ChatGPT, Cursor, custom agents
Web consoleSearch, Explore, Discover, administration
Command lineScripting and operations
REST APIEvery operation, programmatically

Access control that lives below the application.

Conventional search products filter results in application code, where one missed clause leaks data. K-Lake enforces tenant isolation and per-file access with row-level security in the database, so every interface and every new code path inherits the same guarantee and fails closed.

  • Identity-firstEvery caller is a named principal, signed in through your own identity provider over OIDC and OAuth 2.1, or holding a personal access token tied to one user.
  • Source-native permissions honouredNTFS and NFSv4 ACLs and POSIX ownership are captured at crawl time and correlated to your directory identities, then enforced on every query.
  • Memory-safe coreThe services and every source connector are written in Rust, removing whole classes of memory-corruption vulnerabilities by construction.
  • Secrets encrypted before they reach the databaseSource credentials and tokens are sealed with authenticated encryption under a master key that is never stored alongside them. No read API returns a secret.
  • Audited and observableEvery AI query and every open of an original file is recorded with the principal, tenant and files returned. Structured logs and metrics feed your existing SIEM.
  • Continuously scanned, signed imagesA rolling daily security report publishes vulnerability findings across every shipped container image, and release images are cryptographically signed.

Where teams start with K-Lake.

The same index serves very different questions. These are the ones we are asked most.

All use cases, including regulated and air-gapped environments, media archives and multi-tenant providers

White papers for the people who have to sign it off.

Written for architects, security reviewers and data leaders. No registration required.

We build the product. We also build with it.

KDBL started as a specialist enterprise AI consultancy and Microsoft Partner, and that work continues. Whether you are deploying K-Lake or building agentic systems and Copilot integrations of your own, you work with the people who write the software.

Consulting services
  • K-Lake deliveryProof of concept, production rollout, sizing and enablement for your operators.
  • Agentic AI systemsAgents that reason, plan and act across your tools, with human-in-the-loop controls.
  • Copilot and Workspace integrationConnectors and agents for Microsoft 365 Copilot and Google Workspace.
  • AI strategy and governanceReadiness, roadmaps, architecture and responsible-AI risk, from people who ship.

Deployed with partners you already trust.

K-Lake is available through the Azure Marketplace with licensing derived from your plan, and KDBL is a Microsoft Partner. We are building a partner programme for integrators, managed service providers and platform vendors.

Talk to us about partnering

See K-Lake on your own kind of data.

Tell us what you hold and what you want to ask of it. We will walk you through K-Lake on a representative estate and, if it fits, set you up with a 30-day evaluation licence. We reply within one business day.

Evaluation runs on a single VM in under an hour
Your data stays in your environment throughout
No obligation