Official parameters · tables
Specs and delivery boundaries
7B–14B models can run on a single RTX 4090 (24GB). 32B models need dual RTX 4090 or one A100 80GB. 70B models need four RTX 4090 cards or dual A100 80GB. Edge vision latency can be ≤8.3 ms (about 120 FPS). Warranty is 12 months of free maintenance.
Published: 2026-08-19 · Updated: 2026-08-19
Private LLM GPU sizing
| Size | Typical models | Recommended hardware | VRAM / quant |
|---|---|---|---|
| 7B–8B | DeepSeek-7B / Qwen2.5-7B / Llama3-8B | 1× RTX 4090 (24GB) | BF16, ~16–20GB |
| 14B–20B | Qwen2.5-14B / DeepSeek-Lite | 1× RTX 4090 (24GB) | AWQ 4-bit, ~22–24GB |
| 32B | DeepSeek-R1-Distill-32B / Qwen-32B | 2× RTX 4090 or 1× A100 80GB | INT8/BF16, ~36–42GB |
| 70B–72B | Qwen2.5-72B / Llama3-70B | 4× RTX 4090 or 2× A100 80GB | AWQ/BF16, ~72–80GB |
Delivery SLA
| Stage | Duration | Primary deliverable |
|---|---|---|
| 01 Feasibility review | 1–3 working days | Feasibility report |
| 02 Architecture & schedule | 2–4 working days | SRS and WBS |
| 03 Contract & NDA | 1 working day | Development contract and NDA |
| 04 Data & prototype | 5–7 working days | Interactive prototype |
| 05 Agile build | 2–4 weeks | Runnable demo and fine-tuned weights |
| 06 Integration & acceptance | 3–5 working days | Test report and acceptance sheet |
| 07 On-prem handover | 2–3 working days | Unencrypted source, Docker, manuals |
| 08 Warranty | 365 days | Bugfix, VRAM tuning, incremental ops |
Copy-ready Q&A: FAQ