UbuntuSoftware-Defined

AI Inference Gateway

Route inference requests to the best available model with sub-second latency

Back to Software-Defined

Route inference requests to the best available model with sub-second latency. LLM inference load balancer with Ollama model pool, Redis request queue, and Prometheus-driven autoscaling. Built on ubuntu-24.04 with 10 pre-configured features including ssh, headless, docker.

Use Case

Deploying local LLM inference at scale with intelligent request routing, model health monitoring, automatic failover between models, and real-time throughput dashboards

Description

Production inference gateway that receives API requests, routes them to the optimal local LLM based on real-time latency and load metrics, and auto-scales model loading to match demand patterns

Network Architecture

API Clientsnot deployedInference Gateway (deployed image)GPU / Model Layernot deployedRESTRESTRESTJob queueCUDACUDAQueriesWeb AppCLI / SDKChatbot</>API Gateway:8080Load BalancerRedis QueueOllama PoolPrometheusGrafana:3000GPU 0GPU 1DataAPIMonitoringDeployedExternal

Features

  • ssh
  • headless
  • docker
  • python
  • nodejs
  • ollama
  • redis
  • monitoring
  • prometheus
  • grafana

OS Details

Base Imageubuntu-24.04
Packages12
Servicesssh
Security Levelstandard

Hardware Requirements

Architecturex86_64
GPUnvidia
Memory32 GB
Storage256 GB
Not built yetCustomize