Rust
Language
Memory-safe systems code with explicit interfaces.
Rust LLM Serving Engine
Hetero-Paged-Infer concentrates on the core serving path: paged KV cache management, continuous batching, and OpenAI-compatible HTTP APIs.
Memory-safe systems code with explicit interfaces.
Block-based allocation and accounting.
Prefill/decode flow with decode-priority behavior.
Completions and chat endpoints with operational probes.
Engine structure, scheduling model, and memory management design.
Installation, configuration, and local usage.
Core types and HTTP API references.
Docker and production-oriented deployment notes.
Contributing and validation workflow.