LLM Server — On-Device AI API
Run AI models on your phone and serve them as an API endpoint to any app
About LLM Server — On-Device AI API
★ Turn Your Phone Into an AI Server ★
LLM Server lets you run large language models directly on your Android device — and expose them as a local API endpoint that any app, script, or AI agent can connect to.
No cloud. No subscriptions. Your phone becomes the AI backend.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🧠 ON-DEVICE INFERENCE
• Run open-weight LLMs (Llama, Gemma, Phi, Qwen, Mistral and more)
• Optimized for mobile hardware with GGUF quantization
• Full privacy — your data never leaves your device
• Works completely offline once models are downloaded
🔌 BUILT-IN API SERVER
• OpenAI-compatible REST API served from your phone
• Connect any tool that speaks the OpenAI API format
• Stream responses in real-time via Server-Sent Events
• Serve multiple clients on your local network
🤖 PERFECT FOR AI AGENTS
• Let your AI assistants call your phone as an LLM backend
• Hermes Agent, LangChain, AutoGPT — anything with HTTP
• Tailscale/ZeroTier support for secure remote access
• No API keys needed — you own the model
📱 HOW IT WORKS
1. Download a model from the built-in model browser
2. Tap "Start Server" to launch the API endpoint
3. Point any client to http://your-phone-ip:8080/v1/chat/completions
4. That's it — your phone is now an AI server
🔧 TECHNICAL DETAILS
• llama.cpp backend for fast CPU/GPU inference
• Supports GGUF models (Q4_K_M, Q5_K_M, Q8_0 and more)
• OpenAI-compatible /v1/chat/completions endpoint
• Configurable context length, temperature, and sampling
• Background service keeps serving even when app is minimized
• Battery-optimized inference scheduling
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
💡 USE CASES
• Personal AI assistant that runs 100% on your device
• Development & testing — prototype AI features without cloud costs
• AI agent backend — let autonomous agents use your phone as their brain
• Privacy-first AI — sensitive data never touches the cloud
• Offline AI — works on planes, in tunnels, anywhere
🔒 PRIVACY & SECURITY
• All inference happens on-device
• No data collection, no telemetry, no cloud dependency
• API server only accessible on your local network by default
• Open-source model ecosystem — no vendor lock-in
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
LLM Server bridges the gap between powerful open-source AI models and the devices you carry every day. Stop paying for cloud API calls — run your own AI, on your own terms.
Download now and turn your phone into an AI powerhouse.