LLMKube
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API
Categories
Platforms
License
Last Updated
September 07, 2026Current Release
Commit Activity
Related Applications
GoModel
1.1kAI gateway written in Go with a unified OpenAI-compatible API for multiple LLM providers, USD cost...
Garlic-Hub
143Digital signage device and content management system with SMIL playlist support and scheduling
Ollama
180.4kGet up and running with Llama 3.3, DeepSeek-R1, Phi-4, Gemma 3, and other large language models
Open-WebUI
151.3kUser-friendly AI Interface, supports Ollama, OpenAI API
LocalAI
49.0kRun your AI models locally and generate images and audio
Onyx Community Edition
32.0kChat UI that works with any LLM. It comes loaded with advanced features like agents, web search,...