Benchmarks for small (~4B) tool-calling open-weight LLMs on commodity x86 CPU — standard vs TurboQuant KV-cache compression, measuring tool-calling accuracy, throughput, and memory.
-
Updated
May 22, 2026 - Python
8000
Benchmarks for small (~4B) tool-calling open-weight LLMs on commodity x86 CPU — standard vs TurboQuant KV-cache compression, measuring tool-calling accuracy, throughput, and memory.
A FastAPI server for querying Google's Gemma Translate AI models for translations
Local-first CLI, Docker, and Tauri workspace for hardware-aware Kimi K3 inference on consumer systems.
Running Llama 2 and other Open-Source LLMs on CPU Inference Locally for Document Q&A
Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command
⚡ LLM Gateway is a zero-config local AI gateway. Run any GGUF model on your CPU with one command to instantly spin up an OpenAI-compatible API gateway. The ultimate lightweight LLM gateway—no Python, no Docker, no GPU required.
Lightweight web UI for llama.cpp with dynamic model switching, chat history & markdown support. No GPU required. Perfect for local AI development.
Open source lightweight Sanskrit language model optimized for CPU, edge devices, and offline inference.
Compress PyTorch models for edge devices — CPU-only, no GPU, no retraining. One function call.
Dockerized RAG system optimized for CPU inference. No GPU required.
A local model router: a tiny always-on model classifies each request and hot-swaps a specialist in via Ollama only when one is needed. Runs entirely on CPU.
CPU-only FastAPI server for Typhoon ASR Realtime
Real-time facial emotion detection using DeepFace + OpenCV — CPU-only, no GPU required. Frame-skip optimisation for live webcam inference.
Un sistema RAG per chattare con documenti locali usando Foundry e modelli LLM su CPU
Deepixel develops real-time human understanding algorithms from a single RGB camera, with full in-house capability spanning data acquisition, annotation, model training, and platform-optimized deployment—ensuring robust, efficient, and production-ready vision systems.
llama-swap Docker image bundling an ik_llama.cpp server alongside mainline llama.cpp (amd64)
Pocket TTS Setup for Windows. Easy to Use.
CPU-only local coding agents with OpenCode, llama.cpp and interchangeable GGUF models — tested on Intel Ivy Bridge.
CPU-efficient face age and gender models and training pipeline
End-to-end receipt extraction pipeline using LayoutLMv3 + EasyOCR with a Streamlit Human-in-the-Loop review interface. Runs entirely on CPU.
Add a description, image, and links to the cpu-inference topic page so that developers can more easily learn about it.
To associate your repository with the cpu-inference topic, visit your repo's landing page and select "manage topics."