vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm
-
Updated
Oct 11, 2026 - Python
vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm
Freeze Claude Code's prompt prefix so DeepSeek's automatic cache always hits — alignment proxy + coalescing + keepalive, installable as a CC plugin. Measured 64% cheaper on real Claude Code traffic.
Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.
DevWhale —— AI 驱动桌面开发工作台。深度契合Deepseek V4,做了针对性缓存优化。Electron + React + TypeScript,流式 Agent对话、Monaco 编辑器、xterm.js 终端、60+ 文件格式多模态输入、多模型切换。DevWhale — AI desktop dev workbench. Electron + React + TS. Streaming Agent chat,Monaco editor, xterm terminal, 60+ file formats, multi-model support.
Radix trie-based prompt prefix cache optimizer matching longest common prompt token prefixes to eliminate redundant prefill computation and reduce TTFT.
Radix trie-based prompt prefix cache optimizer matching longest common prompt token prefixes to eliminate redundant prefill computation and reduce TTFT.
An experimental KV mobility runtime.
RACS (Remote Agent Context Store): prefix-cache management for production agents. Stability-aware prompt planning, provider-faithful cache directives for Anthropic, OpenAI, Gemini, Bedrock and more, TTL keep-warm scheduling, prefix-drift detection, and hit-ratio and savings analytics. Zero dependencies, TypeScript, edge-ready.
ComfyUI Qwen-Image 2.1 multi-GPU inference: Ulysses sequence parallelism for 1/2/4/8 GPUs, T2I, image editing, INT8/FP8, and distributed prefix cache.
A prefix-cache profit-and-loss layer for DeepSeek coding agents — wrap your client in two lines, see per-request cache HIT/PARTIAL/MISS, the ¥ each prefix-bust wasted, and the reorder that restores the discount.
DeepSeek缓存优化器 v1.1 — Reasonix四支柱 + 语义压缩 (命中率+30%)
Codex + DeepSeek 前缀缓存医生:本地透明代理,稳定请求前缀提升缓存命中率,附网页控制台
End-to-end LLM serving simulator integrating scheduling, prefix caching, tensor allocation, and KV-cache management. 168-run sweep (72 baseline + 96 pressure). Key finding: ChunkedPrefill + LFU cache achieves 41% lower TTFT p95 and 94% prefix hit rate, but hits OOM first under memory pressure.
Hermes plugin: DeepSeek prefix-cache wire shaping (Reasonix-inspired)
Persistent KV prefix caching middleware for llama.cpp and OpenAI-compatible agents. Detects stable system/developer prompts and tools, persists slot snapshots, and restores them after model restarts.
Cost governance for Massive Intelligence (IM) agent orchestration: hard per-request, per-task, and per-day USD budgets with depletion events, pre-flight cost forecasting, prefix-cache break-even planning across sixteen provider profiles, and cheapest-first cascade routing. Zero dependencies, node-free core, deterministic, TypeScript-first.
A lightweight LLM inference and tool-calling agent system with KV cache experiments.
Consumer AI that remembers your whole life — affordable on DeepSeek prefix-cache
ctxfeed is a local MCP project-context backend that shards a whole repo into GLM-5.2's 1M-token window with cache-aware ingest ordering
Event-driven simulator for prefix KV-cache eviction policies in LLM serving systems
To associate your repository with the prefix-cache topic, visit your repo's landing page and select "manage topics."