ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

inference-efficiency

1 post tagged with "inference-efficiency"

July 23, 2026

SALT: Salience-Aware Lexical Trie for Long-Context Compression

Model-agnostic extractive prompt compression framework using sentence-frequency-ordered lexical trie to prevent theme collapse under tight budgets

prompt-compression long-context kv-cache extractive-compression theme-collapse trie-indexing inference-efficiency prefill-optimization

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.