ctxprune: context compression for AI agents

A small token classifier deletes the least useful words from tool outputs, logs, code and documents before an LLM reads them. It is a successor to Microsoft's LLMLingua-2, trained on agent text, and it never splits or re-spaces identifiers.

68.8% vs 53.4%questions an LLM still answers at 50% kept, ctxprune vs LLMLingua-2-large (91.8% uncompressed)
53.4% vs 32.5%same at 33% kept
~6× fasteron CPU than LLMLingua-2-large; int8 ONNX, no torch needed

struck = deleted · red = identifier deleted or mangled (what the LLM no longer sees intact)

Real outputs of darioooooo0o/ctxprune-small (int8 ONNX) and microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank, computed offline with each library's own rate. LLMLingua-2 measures rate in its own tokens and often keeps more than asked. QA numbers: 378 held-out questions answered by an independent LLM, details on the model card.