SUSE AI
SUSE AI: enterprise GenAI platform. Secure on-premises LLM deployments, RAG, Rancher integration. Open source AI for enterprise.

Key Features
- On-Premises GenAI - LLM without sending data to cloud
- Open Source Models - Llama, Mistral, no vendor lock-in
- RAG Ready - Retrieval Augmented Generation
- Kubernetes Native - deployment on Rancher/K8s
- GPU Support - NVIDIA GPU acceleration
Table of Contents
What is SUSE AI?
SUSE AI is an enterprise platform for deploying Generative AI - allows you to run large language models (LLM) on-premises, without sending data to external clouds. Based on open source, integrated with SUSE ecosystem (Rancher, SLES).
Key Functions:
- On-Premises LLM - Llama, Mistral, other models locally
- RAG - Retrieval Augmented Generation with your own data
- Kubernetes Native - deployment on Rancher
- Air-Gap Ready - for regulated industries
What problem does it solve?
flowchart LR
subgraph Cloud AI Problem
A[Company data] --> B[OpenAI/Azure]
B --> C[Data in cloud]
C --> D[Compliance risk]
D --> E[Vendor lock-in]
end
subgraph SUSE AI
F[Company data] --> G[SUSE AI On-Prem]
G --> H[Data stays local]
H --> I[Compliance OK]
I --> J[Open source]
end
style A fill:#6366f1,stroke:#4f46e5,color:#fff
style B fill:#dc2626,stroke:#b91c1c,color:#fff
style C fill:#dc2626,stroke:#b91c1c,color:#fff
style D fill:#dc2626,stroke:#b91c1c,color:#fff
style E fill:#dc2626,stroke:#b91c1c,color:#fff
style F fill:#6366f1,stroke:#4f46e5,color:#fff
style G fill:#22c55e,stroke:#16a34a,color:#fff
style H fill:#22c55e,stroke:#16a34a,color:#fff
style I fill:#22c55e,stroke:#16a34a,color:#fff
style J fill:#22c55e,stroke:#16a34a,color:#fff
Common Problems:
- Company data sent to OpenAI/Azure AI = compliance risk
- Regulations (banking, healthcare) prohibit cloud AI
- Vendor lock-in to OpenAI/Azure
- No control over model and training data
- API call costs in cloud
SUSE AI Architecture
flowchart TD
A[Users/Apps] --> B[SUSE AI API]
B --> C[Inference Engine]
C --> D[LLM Models]
D --> E[Llama 3.x]
D --> F[Mistral]
D --> G[Custom Models]
B --> H[RAG Pipeline]
H --> I[Vector DB]
H --> J[Enterprise Data]
K[Rancher] --> C
L[NVIDIA GPU] --> C
style A fill:#6366f1,stroke:#4f46e5,color:#fff
style B fill:#f59e0b,stroke:#d97706,color:#fff
style C fill:#8b5cf6,stroke:#7c3aed,color:#fff
style D fill:#22c55e,stroke:#16a34a,color:#fff
style H fill:#22c55e,stroke:#16a34a,color:#fff
style K fill:#f59e0b,stroke:#d97706,color:#fff
style L fill:#22c55e,stroke:#16a34a,color:#fff
Key Features
Open Source LLMs
No vendor lock-in
- Llama 3.x (Meta)
- Mistral / Mixtral
- Custom fine-tuned models
- GGUF, safetensors
RAG Pipeline
Your own data
- Document ingestion
- Vector embeddings
- Semantic search
- Context injection
GPU Acceleration
NVIDIA GPUs
- CUDA support
- Multi-GPU inference
- Quantization (4-bit, 8-bit)
- Tensor parallelism
Kubernetes Native
Rancher deployment
- Helm charts
- GPU scheduling
- Auto-scaling
- Rancher integration
Enterprise Security
Compliance ready
- Air-gap deployment
- Data stays on-prem
- RBAC
- Audit logging
API Compatible
OpenAI-compatible
- OpenAI API format
- Drop-in replacement
- SDK compatibility
- Easy migration
Use Cases
Internal Knowledge Base
RAG on company documentation - employees ask AI about procedures, policies, products. Data doesn't leave company.
Code Assistant
LLM for developers - code completion, review, documentation. Source code stays internal.
Customer Support
AI chatbot based on your own product data. First-line support automation.
Document Processing
Contract, invoice, report analysis. Data extraction without sending to cloud.
SUSE AI vs Cloud AI
| Aspect | SUSE AI | OpenAI/Azure AI |
|---|---|---|
| Data location | On-premises | Cloud |
| Compliance | Air-gap ready | Depends |
| Cost model | Infrastructure | Per token |
| Vendor lock-in | Open source | Proprietary |
| Customization | Full control | Limited |
| Latency | Low (local) | Network dependent |
| Models | Open source | OpenAI only |
| Scaling | Your infrastructure | Automatic |
SUSE AI Advantage:
- Data stays on-premises
- Compliance for regulated industries
- Open source models - no vendor lock-in
- Predictable costs (infra vs per-token)
- Full customization and fine-tuning
Who is it for?
SUSE AI MAKES sense when:
- • Data cannot leave company (compliance)
- • Banking, healthcare, government
- • You have GPU infrastructure
- • Want to avoid vendor lock-in
- • Have Rancher/SLES (SUSE stack)
SUSE AI DOESN'T make sense when:
- • Don't have GPU infrastructure
- • Need GPT-4 class models
- • Low volume - cloud AI cheaper
- • No team to maintain
Requirements
| Parameter | Minimum | Recommended |
|---|---|---|
| GPU | NVIDIA RTX 3090 / A10 | NVIDIA A100 / H100 |
| VRAM | 24 GB | 80 GB+ |
| RAM | 64 GB | 256 GB+ |
| Storage | 500 GB SSD | 2 TB NVMe |
| OS | SLES 15 SP5+ / Linux Micro | SLES 15 SP6 |
| Kubernetes | K3s / RKE2 / Rancher | Rancher managed |
Specification
| Parameter | Value |
|---|---|
| Models | Llama 3.x, Mistral, Mixtral, custom |
| API | OpenAI-compatible REST API |
| Deployment | Kubernetes (Rancher, K3s, RKE2) |
| GPU | NVIDIA CUDA |
| Quantization | 4-bit, 8-bit, FP16 |
| RAG | Vector DB integration |
| Security | Air-gap, RBAC, audit |
FAQ
What models does it support? Open source LLMs: Llama 3.x, Mistral, Mixtral, CodeLlama, and others. You can also use your own fine-tuned models.
Do I need GPU? Yes for production. NVIDIA GPUs with minimum 24GB VRAM. For dev/test can use CPU (slower).
Is the API compatible with OpenAI? Yes. SUSE AI offers OpenAI-compatible API - easy migration from cloud AI.
How does RAG work? You import documents, system creates vector embeddings, on question searches relevant fragments and adds to LLM context.
Can models be fine-tuned? Yes. Platform supports fine-tuning on your own data.
Does nFlo deploy SUSE AI? Yes. Deployments on Rancher/K8s, GPU configuration, RAG pipeline setup, application integration.
Inquire about SUSE AI
Contact your product specialist and get a custom quote.

Related Services
Our services supporting the implementation and management of this solution
Strategic AI and GenAI Implementations in Business
AI and Automation
Transform your business with AI. Strategic implementations that deliver measurable ROI.
AIOps - AI for IT Operations
AI and Automation
Stop fighting fires. AIOps predicts problems before they impact business.
AI Chatbots - Assistant Implementations
AI and Automation
A chatbot that actually helps. Not frustrates. AI assistant implementations with 80% resolution rate.
IBM watsonx - Enterprise AI Platform
AI and Automation
AI for business, not for hype. IBM watsonx implementations with ROI from month one.
Related Products
Other solutions you might be interested in
HCL BigFix
HCL
HCL BigFix: unified endpoint management. Patching, compliance, security for 100+ OS. On-prem, cloud, remote - single agent.
HCL Volt MX
HCL
HCL Volt MX: low-code development platform. Multi-experience apps, rapid development, enterprise integration.
HCL Workload Automation
HCL
HCL Workload Automation: enterprise job scheduling and orchestration. Kubernetes-native, cloud-ready, self-service workflows.
IBM Apptio
IBM
IBM Apptio: FinOps and IT cost management platform. Shows how much you spend on IT/cloud, who pays for what, where to save. Cloudability for AWS/Azure/GCP optimization.
Want to Reduce IT Risk and Costs?
Book a free consultation - we respond within 24h
Or download free guide:
Download NIS2 Checklist