Skip to content

Repository files navigation

clef-compactor logo

clef-compactor

Keep the evidence. Cut the noise.

Query-aware RAG context compaction with Cloudflare Clef.
Score an entire retrieval batch against the query, retain the chunks that earn their place, and send a smaller, auditable context to your LLM.

CI PyPI Python 3.10+ Cloudflare Clef License

Get started · Watch the demo · How it works · Benchmarks · Kaggle notebook


Illustration of retrieved documents passing through a relevance gate: useful evidence is retained and irrelevant context is redirected

Demo of clef-compactor ranking retrieved chunks, cutting irrelevant context, and preserving the useful evidence
▶ 12-second product demo · Open video · View landing page

Why clef-compactor?

Retrieval can be noisy. If a retriever returns 12 chunks, your LLM normally reads—and you pay for—all 12, even when several are off-topic. clef-compactor asks Clef whether each chunk is needed for the user's query, ranks the answers, then fills your token budget with the strongest evidence.

What it does Why it matters
Keeps text verbatim Citations and provenance remain intact—nothing is summarized or rewritten.
Explains every cut Dropped chunks are marked irrelevant or budget_exhausted.
Batches requests One Cloudflare call scores up to 64 chunks.
Fits your stack Sync and async clients, CLI, OpenAI-compatible handler, LangChain, and LlamaIndex adapters.

Demo

The video above is embedded as a clickable preview for GitHub compatibility. You can also play it directly here:

It shows five retrieved chunks flowing through the relevance gate: useful passages are retained, unrelated context is removed with a reason, and the final context is 34% smaller.

Quick start

pip install clef-compactor

# PowerShell
$env:CLEF_ACCOUNT_ID = "your_account_id"
$env:CLEF_API_TOKEN = "your_api_token"

# bash / zsh
export CLEF_ACCOUNT_ID=your_account_id
export CLEF_API_TOKEN=your_api_token
from clef_compactor import ClefCompactor

compactor = ClefCompactor()
result = compactor.compact(
    "What is the refund policy?",
    retrieved_chunks,
    token_budget=1_000,
)

context = "\n\n".join(result.kept_texts())  # exact original chunk text
print(result.saved_fraction)                   # e.g. 0.71 = 71% fewer tokens
print(result.dropped[0].drop_reason)           # irrelevant | budget_exhausted

Need a runnable, credential-free tour? Open the example notebook, which uses a canned API response.

Async

from clef_compactor import AsyncClefCompactor

compactor = AsyncClefCompactor(model="clef-flash")
result = await compactor.compact(query, chunks, token_budget=800)
await compactor.aclose()

CLI

clef-compact -q "refund policy" -d "chunk one" "chunk two" --budget 1000
clef-compact -q "..." -d @chunk1.txt @chunk2.txt --json --model clef-flash

How it works

Diagram: Clef scores each retrieved chunk, relevant chunks pass the threshold, and the remainder are cut with an auditable reason
  1. Build a decision for each chunk. The query and a configurable preview of every retrieved chunk are sent as Clef noul questions.
  2. Score the batch. Clef returns P(relevant) for each chunk; batches larger than 64 are split automatically.
  3. Rank and fit the budget. Chunks are sorted by relevance score, then token count and original position; the best ones are greedily retained until the budget is full.
  4. Keep an audit trail. The result contains kept and dropped chunks, scores, token counts, API usage, latency, and a cost estimate.

Integrations

OpenAI-compatible LangChain LlamaIndex
Wrap a FastAPI endpoint and return an OpenAI response. Use a document compressor in your retrieval pipeline. Add a node postprocessor to an existing query engine.
OpenAI-compatible endpoint
from fastapi import FastAPI
from clef_compactor import ClefCompactor
from clef_compactor.compat.openai import handle_chat_completions

app = FastAPI()
compactor = ClefCompactor()

@app.post("/v1/chat/completions")
async def chat_completions(payload: dict):
    return handle_chat_completions(payload, compactor=compactor)

Put chunks in clef.chunks, or supply context as non-user messages. The compacted context is returned in choices[0].message.content; before/after token counts are in usage.

LangChain
pip install "clef-compactor[langchain]"
from clef_compactor.integrations.langchain import ClefDocumentCompressor

compressor = ClefDocumentCompressor(token_budget=800)
compressed = compressor.compress_documents(docs, query="refund policy?")
LlamaIndex
pip install "clef-compactor[llamaindex]"
from clef_compactor.integrations.llamaindex import ClefNodePostprocessor

query_engine = RetrieverQueryEngine(
    retriever=base_retriever,
    node_postprocessors=[ClefNodePostprocessor(token_budget=800)],
)

Benchmarks

Local, open-weights evaluation

clef-flash 9B in float16 on 2× Kaggle T4 GPUs, evaluated over 20 labelled cases (104 chunks). The evaluation data and result are committed in this repository; the GPU run is available in the Kaggle notebook ↗.

Metric Result
Chunk accuracy 0.712
Kept precision 0.771
Relevant recall 0.746
Kept F1 0.758
Context tokens saved 38.4%
Scoring latency (p50 / p95) 1,274 / 1,407 ms
Cost per 1k calls $0.00 (self-hosted weights)

Run the deterministic pipeline validation locally:

python evals/run_eval.py

The replay scorer validates the ranking and budget machinery—not model quality—and scores 0.990 chunk accuracy, 0.992 kept F1, and 34.0% saved tokens. For a live hosted-API run, configure credentials and use python evals/run_eval.py --mode live.

Clef and Laya: the decision-quality trade-off

Published figures from Cloudflare's Introducing Clef post:

Benchmark Clef Clef Flash Laya
BFCL, case exact 98.47 98.76 38.13
ToolRet, nDCG@10 69.19 66.43 12.69
API-Bank accuracy 91.93 93.11 11.41
Median latency 209.3 ms 38.8 ms 5.8 ms
Context window 65,536 65,536 32k (reported)

In short: Laya is the lower-latency local option; Clef gives substantially stronger decision quality; Clef Flash is the pragmatic middle ground. clef-compactor adds one network round trip per 64 chunks.

Configuration

Variable Default Purpose
CLEF_ACCOUNT_ID / CLOUDFLARE_ACCOUNT_ID required Cloudflare account ID
CLEF_API_TOKEN / CLOUDFLARE_API_TOKEN required Token with Workers AI run permission
CLEF_MODEL clef clef or clef-flash
CLEF_BASE_URL Cloudflare v4 API Override for AI Gateway or tests
CLEF_TIMEOUT 60 Per-request timeout in seconds
CLEF_MAX_RETRIES 2 Retries for 408 / 429 / 5xx and network failures
CLEF_LOG_LEVEL WARNING Standard-library logger level

Errors derive from ClefError, so a single except covers authentication, rate limits (including retry_after), server failures, timeouts, network errors, and malformed responses. Retries use exponential backoff with jitter and honour Retry-After.

Limitations

  • Evaluation scope. The local result is from a 9B model on two T4 GPUs—not Cloudflare's hosted endpoint or the larger 27B model. Run --mode live before quoting hosted-model quality.
  • Latency. Model decision latency is 209 ms for Clef and 38.8 ms for Clef Flash, plus network time. If every millisecond matters, see laya-compactor.
  • Estimated token counts. Budgets use cl100k_base as a close proxy rather than Cloudflare's exact tokenizer.
  • Preview scoring. Only the first 512 characters of each chunk are scored by default. Increase chunk_preview_chars if important context is deep in a chunk.
  • Binary relevance. Clef emits P(relevant), not a graded “supporting vs. essential” signal.

Project links

License

Apache-2.0. Clef itself is open source on Hugging Face under the same license.

About

Query-aware RAG context compaction with Cloudflare Clef: keep the evidence, cut noisy retrieval context.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages