Seed AI Infrastructure

System Requirements

Seed AI is self-hosted and runs on infrastructure controlled by your business. The application, locally hosted language models, company documents, knowledge bases, user accounts, and chat history remain within your own server environment.

The exact hardware your business needs depends on the AI models you plan to run, how many employees will have access, how many users may actively use AI at the same time, your context and knowledge requirements, and the response performance you expect.

Required Environment

The core infrastructure
needed to run Seed AI.

A typical production deployment requires a Windows Server environment with dedicated GPU resources, IIS, SQL Server, sufficient system memory and high-speed storage, and secure network access for authorized users.

Windows Server

A supported Microsoft Windows Server environment under the control of your business.

NVIDIA GPU

Production GPU capacity sized for your selected models, expected concurrent users, context requirements, and desired response performance.

Microsoft IIS

IIS installed and configured to host the private Seed AI web application.

SQL Server

Microsoft SQL Server for application, user, model, agent, configuration, knowledge, and chat data.

Business Network

Secure network access for employees, administrators, and any approved remote users.

SSD / NVMe Storage

High-speed storage for AI models, company documents, indexes, application data, logs, exports, and backups.

Server Environment

Runs on a Windows Server in your network.

Seed AI runs as a private business application within a supported Microsoft Windows Server environment.

The server may be dedicated to Seed AI or provided through another supported configuration that gives the application direct access to the required NVIDIA GPU resources.

For production and multi-user deployments, a dedicated GPU-enabled server is recommended for predictable performance, capacity planning, and future expansion.

Recommended Server Environment

Operating System Windows Server 2025
Web Server Microsoft IIS
Application Runtime Required .NET runtime
Supporting Runtime Microsoft Visual C++ components
AI Processing Supported NVIDIA GPU
GPU Drivers Supported NVIDIA drivers
Database Microsoft SQL Server 2025
Storage Enterprise SSD / NVMe recommended
Updates Current Microsoft security updates
GPU Requirements

Plan GPU capacity around your users and AI workload.

GPU requirements for a business deployment depend on much more than whether an AI model can technically fit into GPU memory.

Seed AI is a centralized multi-user platform. The AI model can remain loaded and serve multiple employees, but simultaneous conversations consume additional context resources and share the available GPU processing capacity.

Production systems should therefore include enough GPU memory and processing headroom to maintain useful response times as multiple employees use AI throughout the workday.

Production capacity planning  Users + concurrency + model size + context + workload + performance headroom 
Licensed Users Expected Simultaneous AI Use 7–14B Models 30–34B Models 70B Models
Up to 10 1–3 active users 1 × 48 GB GPU Recommended small-business production starting point with room for context and shared use. 1 × 96 GB GPU Provides substantial model and runtime headroom for a small multi-user environment. 2 × 96 GB GPUs High-memory configuration for larger models with production headroom.
11–20 2–5 active users 1 × 96 GB GPU Recommended where several employees may use AI regularly during the workday. 2 × 96 GB GPUs Provides additional memory and processing capacity for concurrent workloads. 2–4 × 96 GB GPUs Multi-GPU production environment sized according to context and expected usage.
21–50 5–10 active users 2 × 96 GB GPUs Or dedicated inference workers where higher throughput is required. 2–4 × 96 GB GPUs Multi-GPU production environment with capacity for sustained shared use. 4 × 96 GB GPUs or more Workload-dependent multi-GPU or multi-server architecture.
51–100 10–20 active users 2–4 × 96 GB GPUs Multiple inference workers may provide better throughput than one model instance. 4 × 96 GB GPUs or more High-throughput multi-GPU environment. Multiple AI Servers Dedicated inference capacity should be evaluated.
101–250 20–40 active users 4 × 96 GB GPUs or more Multiple inference workers or servers may be preferred. Multiple AI Servers Engineered high-throughput environment. Multiple AI Servers Dedicated enterprise inference architecture recommended.
250+ 40+ active users Engineered Capacity Multiple GPUs and inference servers. Engineered Capacity Dedicated multi-server environment. Engineered Capacity Enterprise inference architecture.

These figures are conservative planning recommendations for production business environments, not guaranteed minimum or maximum capacities. Actual requirements depend on the specific model, quantization, GPU architecture, context length, prompt size, response length, tools, agents, knowledge-base usage, and actual simultaneous demand. Seed Technologies reviews the hardware requirements for each Seed AI deployment before hardware is purchased.

Licensed Users  The number of employees with Seed AI access does not equal the number actively consuming GPU resources at one time. 
Simultaneous Use  Employees generating AI responses at the same time have a major effect on required processing capacity and response speed. 
Model Size  Larger models require substantially more GPU memory and processing capacity. 
Context Length  Long conversations, large prompts, and retrieved documents increase memory requirements for active sessions. 
GPU Performance  VRAM is only part of the equation. GPU compute capability and memory bandwidth affect response speed under load. 
AI Workload  Agents, tools, vision, embeddings, reranking, document retrieval, and multiple models may require additional resources. 
Model Planning

Smaller models and larger models serve different business needs.

The largest available model is not necessarily the best model for every user or task. A production environment should balance AI capability with response speed, simultaneous usage, and infrastructure requirements.

Smaller Models Speed and Efficiency

Smaller models require less GPU memory and processing capacity and can often support more simultaneous work on the same server infrastructure.

Lower GPU requirements Faster model loading Higher potential throughput Good for focused business tasks More capacity for simultaneous users More models may fit on available hardware
Larger Models Reasoning and Capability

Larger models generally provide stronger reasoning and broader capabilities, but require substantially more GPU memory and processing capacity.

Stronger general reasoning Better handling of complex instructions Larger GPU memory requirements Greater processing per response Lower concurrency on equivalent hardware May require multiple GPUs
GPU 1 96 GB VRAM

GPU 2 96 GB VRAM
Example Installed GPU Memory 192 GB VRAM
Multiple GPUs

Scale beyond a single GPU when the workload requires it.

Seed AI can use supported multi-GPU configurations when a model is too large for one GPU or when a production environment requires greater processing capacity.

Large models can be distributed across multiple GPUs, allowing the inference environment to use more installed GPU memory than is available on a single card.

For larger multi-user installations, multiple GPUs may also be used to increase overall throughput. Depending on the workload, the architecture may use one model across several GPUs, multiple inference workers, dedicated GPUs for different models, or multiple AI servers.

Support models larger than one GPU Increase installed GPU memory Increase processing capacity Support larger simultaneous workloads Use matching GPUs when practical Provide adequate server power Confirm cooling requirements Verify chassis and slot compatibility Provide sufficient PCIe capacity Install supported NVIDIA drivers
Memory and Processing

The GPU is important,
but it is not the entire server.

System memory and CPU resources support Windows Server, IIS, SQL Server, document processing, knowledge-base indexing, model management, user activity, tools, agents, and background services.

64 GB RAM Testing / Limited Use

Appropriate for development, evaluation, smaller models, or limited production environments with light document and simultaneous-user workloads.

128 GB RAM Small Business Starting Point

A practical production starting point for smaller multi-user deployments using knowledge bases, document processing, and locally hosted AI models.

256 GB RAM Recommended Multi-User Setup

Recommended for larger models, more simultaneous users, substantial knowledge bases, document processing, agents, and multiple AI services.

512 GB+ Advanced Deployment

Appropriate for large-scale models, heavy document workloads, high concurrency, multiple inference processes, or larger enterprise installations.

Processor guidance

A modern server-class processor with multiple CPU cores is recommended. Additional CPU capacity may be needed for document processing, indexing, database activity, AI tools, background agents, file processing, and numerous simultaneous users.

Storage Requirements

Plan storage for more than the application itself.

GGUF model files can be large, and Seed AI also requires storage for company documents, knowledge-base indexes, chat history, application data, system logs, exports, and backups.

Enterprise SSD or NVMe storage is strongly recommended. Faster storage improves model loading, document processing, indexing, backups, and general application performance.

Number of AI models Size of stored GGUF models Uploaded document volume Knowledge-base indexes Chat history retention System log retention Export requirements Backup and recovery capacity
Models
Documents
Application Data
Backups
GGUF Models Often the largest individual files
Documents Company files and knowledge-base content
Application Data Chat history, logs, indexes, and configuration
Backups Database, documents, application, and configuration
   SQL
Users Permissions Models Agents Knowledge Bases Chat History Settings Usage Data
Microsoft SQL Server

Your Database Server. Your data control.

Seed AI uses Microsoft SQL Server for application, configuration, user, AI management, and usage data.

SQL Server may be installed on the Seed AI server or provided through another approved SQL Server environment managed by your company.

User accounts Roles and permissions Model configuration AI agent settings Prompt templates Knowledge-base records Chat sessions and history Application settings


Your IT team should include the Seed AI database in its standard backup, maintenance, security, and disaster recovery procedures.

Network Access

Keep access private while supporting the people who need it.

Seed AI is hosted through Microsoft IIS and accessed from a standard web browser. Your IT team controls which users and networks are permitted to reach the platform.

01

Internal Network

Make Seed AI available to authorized users connected to your company's internal network.

02

Company VPN

Allow approved remote employees to connect through your existing secure VPN.

03

Secure Remote Access

Use another approved remote-access configuration controlled by your company's IT team.

Internet Access

Local models do not require a public AI provider.

Seed AI can run locally installed GGUF models without sending your prompts to an outside AI service.

 Internet access may still be used for:  Downloading approved GGUF models Installing Windows and driver updates Downloading Seed AI updates Connecting optional API tools Using approved third-party integrations Accessing external services through an agent
Multi-User Capacity

User licenses and AI capacity are different considerations.

Seed AI is designed as a centralized multi-user business application. Ten employees, twenty employees, or hundreds of employees can have individual accounts while sharing privately hosted AI infrastructure.

The important hardware consideration is how many users may actively ask the AI to perform work at approximately the same time and how demanding those requests are.

Licensed Users Controls how many employees are authorized to use the Seed AI platform.
Simultaneous Users Determines how many AI requests may compete for processing capacity at the same time.
Shared Models A model can remain loaded once and serve requests from multiple authorized users.
Active Context Each active AI sequence consumes additional resources based on its prompt, retrieved information, and conversation length.
GPU Throughput Simultaneous requests share GPU processing resources, affecting how quickly responses are generated.
Growth Planning Production hardware should include capacity for increased AI adoption after deployment.
Capacity Planning

More users do not require a separate GPU for every person.

Seed AI centralizes AI inference on your server environment. Employees access the same private AI infrastructure through the Seed AI application while the server manages the underlying AI workloads.

User Licensing Who Has Access

User licensing controls who can log into Seed AI. The total number of licensed employees does not directly determine the number of GPUs required.

Individual employee accounts Role-based permissions Private chat histories Shared AI infrastructure Shared organizational resources
Server Capacity How Much AI Can Run at Once

Server capacity determines how well the environment handles simultaneous AI work. Active usage is generally more important than the total number of accounts.

Concurrent requests Model size and architecture Context requirements Response speed expectations Agent, tool, and document workloads
Security and Backups

Manage Seed AI like any other business-critical system.

Because Seed AI is installed inside your environment, your business remains responsible for its normal server, network, authentication, backup, security, and recovery procedures.

Restrict server access Apply Windows security updates Maintain supported GPU drivers Use HTTPS and valid SSL certificates Enforce strong user authentication Limit administrative permissions Back up the SQL Server database Back up documents and configuration files Protect model and backup directories Monitor available storage Review system logs Test disaster recovery procedures
Before You Purchase Hardware

You do not have to size the server on your own.

Tell us how your organization plans to use Seed AI and we can recommend GPU, memory, storage, and server specifications based on your users, models, workload, performance expectations, and growth plans.

 Book a Consultation 

Information that helps us size your server:

How many employees will have Seed AI accounts? How many users may actively use AI at once? Which AI models do you want to run? What size models do you expect to use? What response speed do you expect? What context lengths will users require? Will you use private knowledge bases? How many documents will be uploaded? Will document search or RAG be heavily used? Do you need one model or several? Will vision or multimodal models be used? Will agents perform automated or background work? Will users run tools and external integrations? Will multiple departments use different models? Do you already own server hardware? How much future AI adoption should we plan for?
Plan the Right Environment

Start with how your business will use AI. Build the server around the workload.

We can help your IT team select the GPU, system memory, storage, and server configuration required for your Seed AI deployment based on your models, licensed users, expected simultaneous usage, performance requirements, and future growth.