ARTICLE 1. AVAILABLE GPU CATALOG
1.1. High-end GPU line – enterprise AI
| GPU model | VRAM | AI performance (TFLOPS) | Main applications |
| NVIDIA® B300 SXM6 | 192 GB HBM3e | ~2.250 (FP8) | LLM Training, Frontier AI |
| NVIDIA DGX B200 (8x B200) | 1.440 GB | ~18.000 (FP8) | AI Foundation Model |
| NVIDIA DGX H200 (8x H200) | 1.128 GB HBM3e | ~16.000 (FP8) | LLM Fine-tuning, Inference |
| NVIDIA DGX H100 (8x H100) | 640 GB HBM3 | ~8.000 (FP8) | Deep Learning Training |
| NVIDIA DGX A100 (8x A100) | 320 GB HBM2e | ~5.000 (TF32) | AI/ML Research |
| NVIDIA A100 PCIe | 80 GB HBM2e | ~312 (TF32) | ML Training, Inference |
1.2. Mid-range GPU line - professional and HPC
| GPU model | VRAM | Capabilities | Applications |
| NVIDIA L40S | 48 GB GDDR6 | ~733 TFLOPS (FP8) | AI Inference, Rendering |
| NVIDIA L40 | 48 GB GDDR6 | ~362 TFLOPS (FP32) | 3D Rendering, Viz |
| NVIDIA RTX 5090 | 32 GB GDDR7 | ~838 TFLOPS (FP16) | Creative AI, Rendering |
| NVIDIA RTX 4090 | 24 GB GDDR6X | ~330 TFLOPS (FP16) | AI Dev, Gaming, Render |
| NVIDIA RTX PRO 6000 | 96 GB GDDR7 | ~~4.000 (FP8) | Professional AI/Viz |
| NVIDIA RTX 6000 Ada | 48 GB GDDR6 | ~1.457 TFLOPS (FP8) | Professional Workstation |
| NVIDIA A10G | 24 GB GDDR6 | ~250 TFLOPS (FP32) | Inference, Graphics |
| NVIDIA A30 | 24 GB HBM2 | ~330 TFLOPS (TF32) | AI Inference, HPC |
| NVIDIA A40 | 48 GB GDDR6 | ~299 TFLOPS (TF32) | Professional AI/Viz |
| NVIDIA A5000 | 24 GB GDDR6 | ~222 TFLOPS (TF32) | AI Design, Animation |
| NVIDIA A6000 | 48 GB GDDR6 | ~309 TFLOPS (TF32) | High-end Workstation |
| NVIDIA V100 | 32 GB HBM2 | ~125 TFLOPS (FP16) | ML Training (Legacy) |
| NVIDIA A16 | 64 GB GDDR6 | ~1,042 (FP32 total) | VDI, Remote Workstation |
| NVIDIA A4000 | 16 GB GDDR6 | ~153 TFLOPS (TF32) | AI Dev, CAD/CAM |
| NVIDIA RTX 4000 Ada | 20 GB GDDR6 | ~192 TFLOPS (TF32) | Compact Workstation AI |
| NVIDIA L4 | 24 GB GDDR6 | ~485 TFLOPS (FP8) | Video AI, Edge Inference |
ARTICLE 2. GPU RENTAL OPTIONS
2.1. On-demand (hourly)
- Billed by the hour actually used, with no long-term commitment
- Best for: testing, development, irregular workloads
- Launched within 5–15 minutes of the request
- Billed at month end, or when usage reaches the invoicing threshold
2.2. Reserved instance (1, 3, or 12 months)
- Save 20–60% versus on-demand when reserved in advance
- Guarantees GPU resources are available on schedule
- Pay in full up front or monthly
- Best for: steady workloads, long-running model training
2.3. Spot Instance
- 50–80% cheaper than on-demand, using idle capacity
- May be reclaimed with 2 minutes' notice when the resources are needed
- Suited to: batch processing, training with checkpoints and rendering
- Uptime SLA 99.5%; not suitable for workloads requiring continuous availability
- The entire physical server is allocated to the Customer, with nothing shared
- Maximum performance with no “noisy neighbor” effect
- Suited to: HPC clusters and training workflows that require low latency
- Minimum commitment of 1 month
ARTICLE 3. ENVIRONMENT TECHNICAL CONFIGURATION
3.1. Supported operating systems and software
- OS: Ubuntu 20.04/22.04 LTS, CentOS 7/8, RHEL 8/9, Windows Server 2019/2022/2025 (Microsoft SPLA)
- CUDA Toolkit: CUDA 11.x and 12.x (preinstalled, or choose a version)
- Deep learning frameworks: PyTorch, TensorFlow, JAX, PaddlePaddle (images ready to use)
- Container: Docker, NVIDIA Container Toolkit, Singularity/Apptainer
- Kubernetes: NVIDIA GPU Operator, K8s GPU plugin
- Networking: InfiniBand HDR (200 Gb/s) cho multi-GPU cluster, NVLink cho DGX
3.2. Included storage
- Local NVMe SSD: 1–32 TB depending on configuration
- Shared storage: NFS, Lustre, GPFS cho cluster
- Integrated object storage: S3-compatible for large datasets
ARTICLE 4. ACCEPTABLE USE POLICY
4.1. Permitted
- AI/ML: model training, inference, fine-tuning and reinforcement learning
- HPC: scientific simulation, CFD, molecular dynamics and climate modeling
- Rendering: 3D rendering, VFX, real-time visualization
- Video AI: Transcoding, upscaling, object detection
- Academic research and education
4.2. Strictly prohibited
- Cryptocurrency mining – detected immediately, resulting in account lockout
- Cyberattacks, brute force, password cracking
- Training AI models to generate CSAM or other harmful content
- Defeating security systems or exploiting third-party vulnerabilities
- Sharing a GPU account with a third party without a contract
ARTICLE 5. SLA AND SERVICE CREDITS
| Type | SLA Uptime | Maximum downtime per month | Compensation for breach |
| On-demand | 99,9% | 43.8 minutes | 10% of the fee for affected hours |
| Reserved 1 month | 99,9% | 43.8 minutes | 15% of the monthly fee |
| Reserved 12 months | 99,95% | 21.9 minutes | 20% of monthly fee |
| Bare Metal | 99,95% | 21.9 minutes | 25% of the monthly fee |
| Spot Instance | 99,5% | 3.6 hours | No SLA |
ARTICLE 6. PAYMENT
- On-demand: invoiced at month end, or when usage reaches VND 5 million
- Reserved: paid 100% in advance or monthly (the first month on signing)
- Spot: billed per second of actual use, invoiced at month end
- Credit limit exceeded: the instance is suspended automatically 24 hours after the warning
- Refunds: no refund for committed Reserved capacity; the unused portion of On-demand is refunded where the Company breaches the SLA
Revision history
Current versionby HiTechCloud
Updatedby HiTechCloud
Updatedby HiTechCloud