Windows Server
A supported Microsoft Windows Server environment under the control of your business.
Seed AI is self-hosted and runs on infrastructure controlled by your business. The application, locally hosted language models, company documents, knowledge bases, user accounts, and chat history remain within your own server environment.
The exact hardware your business needs depends on the AI models you plan to run, how many employees will have access, how many users may actively use AI at the same time, your context and knowledge requirements, and the response performance you expect.
A typical production deployment requires a Windows Server environment with dedicated GPU resources, IIS, SQL Server, sufficient system memory and high-speed storage, and secure network access for authorized users.
A supported Microsoft Windows Server environment under the control of your business.
Production GPU capacity sized for your selected models, expected concurrent users, context requirements, and desired response performance.
IIS installed and configured to host the private Seed AI web application.
Microsoft SQL Server for application, user, model, agent, configuration, knowledge, and chat data.
Secure network access for employees, administrators, and any approved remote users.
High-speed storage for AI models, company documents, indexes, application data, logs, exports, and backups.
Seed AI runs as a private business application within a supported Microsoft Windows Server environment.
The server may be dedicated to Seed AI or provided through another supported configuration that gives the application direct access to the required NVIDIA GPU resources.
For production and multi-user deployments, a dedicated GPU-enabled server is recommended for predictable performance, capacity planning, and future expansion.
GPU requirements for a business deployment depend on much more than whether an AI model can technically fit into GPU memory.
Seed AI is a centralized multi-user platform. The AI model can remain loaded and serve multiple employees, but simultaneous conversations consume additional context resources and share the available GPU processing capacity.
Production systems should therefore include enough GPU memory and processing headroom to maintain useful response times as multiple employees use AI throughout the workday.
| Licensed Users | Expected Simultaneous AI Use | 7–14B Models | 30–34B Models | 70B Models |
|---|---|---|---|---|
| Up to 10 | 1–3 active users | 1 × 48 GB GPU Recommended small-business production starting point with room for context and shared use. | 1 × 96 GB GPU Provides substantial model and runtime headroom for a small multi-user environment. | 2 × 96 GB GPUs High-memory configuration for larger models with production headroom. |
| 11–20 | 2–5 active users | 1 × 96 GB GPU Recommended where several employees may use AI regularly during the workday. | 2 × 96 GB GPUs Provides additional memory and processing capacity for concurrent workloads. | 2–4 × 96 GB GPUs Multi-GPU production environment sized according to context and expected usage. |
| 21–50 | 5–10 active users | 2 × 96 GB GPUs Or dedicated inference workers where higher throughput is required. | 2–4 × 96 GB GPUs Multi-GPU production environment with capacity for sustained shared use. | 4 × 96 GB GPUs or more Workload-dependent multi-GPU or multi-server architecture. |
| 51–100 | 10–20 active users | 2–4 × 96 GB GPUs Multiple inference workers may provide better throughput than one model instance. | 4 × 96 GB GPUs or more High-throughput multi-GPU environment. | Multiple AI Servers Dedicated inference capacity should be evaluated. |
| 101–250 | 20–40 active users | 4 × 96 GB GPUs or more Multiple inference workers or servers may be preferred. | Multiple AI Servers Engineered high-throughput environment. | Multiple AI Servers Dedicated enterprise inference architecture recommended. |
| 250+ | 40+ active users | Engineered Capacity Multiple GPUs and inference servers. | Engineered Capacity Dedicated multi-server environment. | Engineered Capacity Enterprise inference architecture. |
These figures are conservative planning recommendations for production business environments, not guaranteed minimum or maximum capacities. Actual requirements depend on the specific model, quantization, GPU architecture, context length, prompt size, response length, tools, agents, knowledge-base usage, and actual simultaneous demand. Seed Technologies reviews the hardware requirements for each Seed AI deployment before hardware is purchased.
The largest available model is not necessarily the best model for every user or task. A production environment should balance AI capability with response speed, simultaneous usage, and infrastructure requirements.
Smaller models require less GPU memory and processing capacity and can often support more simultaneous work on the same server infrastructure.
Larger models generally provide stronger reasoning and broader capabilities, but require substantially more GPU memory and processing capacity.
Seed AI can use supported multi-GPU configurations when a model is too large for one GPU or when a production environment requires greater processing capacity.
Large models can be distributed across multiple GPUs, allowing the inference environment to use more installed GPU memory than is available on a single card.
For larger multi-user installations, multiple GPUs may also be used to increase overall throughput. Depending on the workload, the architecture may use one model across several GPUs, multiple inference workers, dedicated GPUs for different models, or multiple AI servers.
System memory and CPU resources support Windows Server, IIS, SQL Server, document processing, knowledge-base indexing, model management, user activity, tools, agents, and background services.
Appropriate for development, evaluation, smaller models, or limited production environments with light document and simultaneous-user workloads.
A practical production starting point for smaller multi-user deployments using knowledge bases, document processing, and locally hosted AI models.
Recommended for larger models, more simultaneous users, substantial knowledge bases, document processing, agents, and multiple AI services.
Appropriate for large-scale models, heavy document workloads, high concurrency, multiple inference processes, or larger enterprise installations.
A modern server-class processor with multiple CPU cores is recommended. Additional CPU capacity may be needed for document processing, indexing, database activity, AI tools, background agents, file processing, and numerous simultaneous users.
GGUF model files can be large, and Seed AI also requires storage for company documents, knowledge-base indexes, chat history, application data, system logs, exports, and backups.
Enterprise SSD or NVMe storage is strongly recommended. Faster storage improves model loading, document processing, indexing, backups, and general application performance.
Seed AI uses Microsoft SQL Server for application, configuration, user, AI management, and usage data.
SQL Server may be installed on the Seed AI server or provided through another approved SQL Server environment managed by your company.
Your IT team should include the Seed AI database in its
standard backup, maintenance, security, and disaster
recovery procedures.
Seed AI is hosted through Microsoft IIS and accessed from a standard web browser. Your IT team controls which users and networks are permitted to reach the platform.
Make Seed AI available to authorized users connected to your company's internal network.
Allow approved remote employees to connect through your existing secure VPN.
Use another approved remote-access configuration controlled by your company's IT team.
Seed AI can run locally installed GGUF models without sending your prompts to an outside AI service.
Seed AI is designed as a centralized multi-user business application. Ten employees, twenty employees, or hundreds of employees can have individual accounts while sharing privately hosted AI infrastructure.
The important hardware consideration is how many users may actively ask the AI to perform work at approximately the same time and how demanding those requests are.
Seed AI centralizes AI inference on your server environment. Employees access the same private AI infrastructure through the Seed AI application while the server manages the underlying AI workloads.
User licensing controls who can log into Seed AI. The total number of licensed employees does not directly determine the number of GPUs required.
Server capacity determines how well the environment handles simultaneous AI work. Active usage is generally more important than the total number of accounts.
Because Seed AI is installed inside your environment, your business remains responsible for its normal server, network, authentication, backup, security, and recovery procedures.
Tell us how your organization plans to use Seed AI and we can recommend GPU, memory, storage, and server specifications based on your users, models, workload, performance expectations, and growth plans.
Book a ConsultationWe can help your IT team select the GPU, system memory, storage, and server configuration required for your Seed AI deployment based on your models, licensed users, expected simultaneous usage, performance requirements, and future growth.