LLMKube
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API
Categories
Platforms
License
Last Updated
July 24, 2026Current Release
Commit Activity
Related Applications
GoModel
1.0kAI gateway written in Go with a unified OpenAI-compatible API for multiple LLM providers, USD cost...
Hiccup
196Beautiful static homepage to get to your links and services quickly. It has built-in search,...
Ollama
176.8kGet up and running with Llama 3.3, DeepSeek-R1, Phi-4, Gemma 3, and other large language models
Open-WebUI
146.6kUser-friendly AI Interface, supports Ollama, OpenAI API
LocalAI
47.8kRun your AI models locally and generate images and audio
Onyx Community Edition
31.1kChat UI that works with any LLM. It comes loaded with advanced features like agents, web search,...