Home / Browse / LLMKube

LLMKube

177 stars

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API

Categories

Generative Artificial Intelligence (GenAI)

Platforms

Go
Docker
K8S

License

Apache, Version 2.0

Last Updated

July 24, 2026

Current Release

v0.9.11
Released July 24, 2026

Commit Activity

Loading graph...

Related Applications

AI gateway written in Go with a unified OpenAI-compatible API for multiple LLM providers, USD cost...

Generative Artificial Intelligence (GenAI) Software Development - API Management

Beautiful static homepage to get to your links and services quickly. It has built-in search,...

Personal Dashboards

Ollama

176.8k

Get up and running with Llama 3.3, DeepSeek-R1, Phi-4, Gemma 3, and other large language models

Generative Artificial Intelligence (GenAI)

User-friendly AI Interface, supports Ollama, OpenAI API

Generative Artificial Intelligence (GenAI)

LocalAI

47.8k

Run your AI models locally and generate images and audio

Generative Artificial Intelligence (GenAI)

Chat UI that works with any LLM. It comes loaded with advanced features like agents, web search,...

Generative Artificial Intelligence (GenAI)