Home / Browse / LLMKube

LLMKube

208 stars

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API

Categories

Generative Artificial Intelligence (GenAI)

Platforms

Go
Docker
K8S

License

Apache, Version 2.0

Last Updated

September 07, 2026

Current Release

v0.9.25
Released September 05, 2026

Commit Activity

Loading graph...

Related Applications

AI gateway written in Go with a unified OpenAI-compatible API for multiple LLM providers, USD cost...

Generative Artificial Intelligence (GenAI) Software Development - API Management

Digital signage device and content management system with SMIL playlist support and scheduling

Miscellaneous

Ollama

180.4k

Get up and running with Llama 3.3, DeepSeek-R1, Phi-4, Gemma 3, and other large language models

Generative Artificial Intelligence (GenAI)

User-friendly AI Interface, supports Ollama, OpenAI API

Generative Artificial Intelligence (GenAI)

LocalAI

49.0k

Run your AI models locally and generate images and audio

Generative Artificial Intelligence (GenAI)

Chat UI that works with any LLM. It comes loaded with advanced features like agents, web search,...

Generative Artificial Intelligence (GenAI)