HiTechCloud AI Platform

Model as a Service (vMaaS)

A platform that delivers AI models as APIs, so businesses can integrate AI quickly, control costs and operate securely on HiTechCloud infrastructure in Vietnam.

What is vMaaS?

Ready-to-use AI models over an API for any business application.

HiTechCloud vMaaS lets development teams call models over an API instead of running complex inference infrastructure themselves, shortening the time it takes to ship AI in a product.

The platform supports chatbots, internal assistants, document processing, content generation, semantic search, RAG, image generation, data analysis and many other enterprise AI automation scenarios.

Service pricing

vMaaS plans that flex with your token usage.

Prices exclude VAT (where applicable). Minimum term: 1 month. Payment: in advance.

You can choose a ready-made token plan or contact HiTechCloud for a custom configuration.

vMaaS_5M

30.000đ

/ month

Starter Dev

5M Tokens / month
8KContent length
8KMax output
UnlimitedRequests/minute
1 monthMinimum term
  • A low entry cost, suited to starting AI experiments without investing in GPU/NPU infrastructure.
  • An OpenAI-compatible API, quick to integrate into existing applications, chatbots or workflows.
  • Access to the platform's full AI model catalog, including LLM, embedding, reranker, vision, and image generation models.
vMaaS_10M

54.000đ

/ month

Freelance Dev

10M Tokens / month
8KContent length
8KMax output
UnlimitedRequests/minute
1 monthMinimum term
  • Capacity suited to developing, testing and regularly tuning AI applications.
  • Supports many different tasks on a single API platform.
  • HiTechCloud infrastructure runs in Vietnam, supporting data security and reliability in production deployments.
vMaaS_100M

480.000đ

/ month

ISV

100M Tokens / month
8KContent length
8KMax output
UnlimitedRequests/minute
1 monthMinimum term
  • A suitable platform for building and delivering AI solutions to multiple end customers.
  • Scale as needed with token add-ons, plus a roadmap for adding new models to the platform.
  • Suited to ISVs, AI agencies and AI solution developers.

Applies to all plans

  • No distinction between input/output tokens
  • No token-per-minute or API-request-per-minute limits
  • Content length up to 8K
  • Max output up to 8K
  • Access the catalog of models available on the platform
Add-on options

Expand capacity or design a custom configuration.

Private consultation

Customize (as needed)

For large enterprises and corporate groups with high token volumes, dedicated configurations, specific security requirements, or BYOM deployments.

Block 1.000.000 token

Add-on Token

6,000 VND/1M tokens

An add-on for customers who need capacity beyond the limits of their base plan. Sold in blocks of 1,000,000 tokens.

Open platform

One API, multiple AI model families for each use case.

LLM Embedding Reranker Image Generation Vision-Language Model BYOM
Key capabilities

Common AI task categories on HiTechCloud vMaaS.

01

Language processing

Chatbots, internal assistants, summarization, document analysis, translation, classification and multi-step reasoning.

02

Image & multimodal

Generate images from text descriptions, recognize and caption images, and process scanned documents, invoices and forms.

03

Embedding & Reranker

Data vectorization, semantic search, RAG, knowledge base Q&A and relevance tuning of results.

04

API & BYOM

Integrate through an OpenAI-compatible REST API/SDK, with support for custom model routing and Bring Your Own Model.

Benefits

Why businesses choose HiTechCloud vMaaS.

Optimized for AI workloads in Vietnam

Supports AI models suited to Vietnamese-language context and to real deployment needs in Vietnam, with the ability to scale by sector — BFSI, retail, logistics, healthcare and the public sector.

Cost-optimized and simple to deploy

Integrate through an OpenAI-compatible API, with no need to invest in GPU/NPU infrastructure or your own MLOps team. The platform is optimized for inference and for operating at enterprise scale.

Data security and operations in Vietnam

Deployed on HiTechCloud's infrastructure in Vietnam, which supports data sovereignty, in-country data residency and information security requirements. Customer data is not used to train third-party models.

Extended AI ecosystem

vMaaS is being developed into a broader AI ecosystem with services such as MLOps, vector database, Knowledge Base-as-a-Service and Agent Builder.

One platform, many models

Connect to multiple families of AI models through a single API standard, covering LLMs, image generation, embeddings, rerankers, and vision-language models.

Suitable from experimentation through to large-scale deployment

Flexible token packages from Starter Dev, Freelance Dev, SME and ISV through to Customize, suited to PoC work, application development and enterprise deployment.

Features

A complete toolset to develop, manage and operate AI.

AI applications

An interface for interacting with and visually inspecting results from common models such as LLM, text-to-image and embedding.

Development tools

An OpenAI-compatible REST API and SDK make integration into existing applications quick, backed by technical documentation, code samples and a sandbox environment for developers.

Model management

Manage the AI model catalog, configure usage limits, and set the direction for supporting an organization's own custom models (BYOM - Bring Your Own Model).

Account & usage management

Supports API key management, resource usage tracking, token consumption reporting, and service billing and credit management.

Support & operations

Documentation, FAQs and technical support from the HiTechCloud team keep AI deployment and operations running smoothly.

AI model catalog

A collection of LLM, embedding, reranker, image generation and vision-language models.

A broad collection of AI models covering a range of enterprise use cases, from conversation and RAG to image generation and multimodal document processing.

LLM (Large Language Models) group

Large language models for conversation, reasoning, summarization, coding, RAG, and business process automation.

Qwen/Qwen3-32B

32B Parameters Reasoning & Agentic

Suited to multi-step reasoning, in-depth document analysis, coding assistance, advanced RAG, and tool-calling process automation.

Qwen/Qwen3-14B

14B Parameters Balanced Chat

Balances performance and cost — suited to chatbots, text analysis, summarization, reasoning and business tasks that require high accuracy.

Meta-llama/Llama-3.1-8B

8B Parameters Multilingual Instruct

A multilingual, long-context model suited to summarization, question answering, tool use, reasoning and coding assistance.

Aisingapore/Llama-SEA-LION-v3.5-8B-R

8B Parameters Regional Optimization

Tuned for Southeast Asian languages and culture, with Vietnamese and multilingual conversation support.

Qwen/Qwen3-4B-Instruct-2507

4B Parameters Edge & Fast Inference

A compact, efficient model for instruction following, logical reasoning, reading comprehension, mathematics, coding and tool calling.

Embedding, reranker & multimodal models

A model set for image generation, data vectorization, search optimization, and vision-language tasks.

Stabilityai/sdxl-turbo

Image Gen Fast Text-to-Image

Generate images from text with fast response times.

BAAI/bge-m3

Embedding Semantic Vectorization

Embeddings for semantic search, RAG, knowledge base Q&A and duplicate content detection.

BAAI/bge-reranker-v2-m3

Reranker Search Optimization

Re-rank results by relevance in the RAG pipeline.

Qwen/Qwen 2.5-VL-7B

Vision LLM Multimodal AI

Vision-language models for image recognition, image captioning, visual reasoning and processing scanned documents, invoices and forms.

Enterprise use cases

Apply vMaaS to your products and operational processes.

AI Chatbot CSKH

A virtual assistant that answers customer questions automatically, 24/7, on your website, app, Zalo OA or other support channels.

Recommended models: Qwen3-32B / Qwen3-14B

Q&A on your internal knowledge base

Build a Q&A system on top of processes, policies, product documentation and operational guides; supports accurate search and answers with source citations.

Recommended models: Qwen3-4B-Instruct / Qwen3-32B / BGE-M3 / BGE-Reranker-V2-M3

Document summarization & analysis

Summarize contracts, financial reports, meeting minutes, and long documents, with support for extracting key points, comparing documents, and synthesizing insights.

Recommended models: Qwen3-14B / Llama3.1-8B

Coding assistant

Code generation, code suggestions, code review, defect detection and IDE integration through an OpenAI-compatible API.

Recommended models: Qwen3-32B / Qwen3-14B

Marketing content generation

Assists with blog posts, social content, marketing email, product descriptions and content variants in different styles.

Recommended models: Qwen3-32B / Qwen3-14B / Llama3.1-8B

Process automation (RPA + AI)

Classify email, extract information from forms and documents, and call tools or functions to complete internal processes.

Recommended models: Qwen3-14B / Qwen3-4B-Instruct
curl https://api.hitechcloud.vn/v1/chat/completions \
  -H "Authorization: Bearer <API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen3-14b","messages":[{"role":"user","content":"Xin chào vMaaS"}]}'
AI Platform Playground

Test models before integrating them into applications.

  • Compare different models on latency, accuracy, cost and support.
  • Submit an interactive trial request and evaluate the output before moving to production.
  • Tune parameters such as temperature, max tokens, and top_p to find the right configuration for each use case.
Build with the API

An OpenAI-compatible API, integrated quickly with an SDK or cURL.

After testing in the Playground, businesses can access the models through HiTechCloud vMaaS to run model inference without additional integration work.

Step 01

Trial

Explore models in the AI Platform Playground and pick the right one based on real results.

Step 02

Create an API key

Create access keys, set application permissions and apply usage limits per environment.

Step 03

SDK integration

Connect via REST API, Python or JavaScript SDK, or call directly from the terminal with cURL.

Step 04

Usage tracking

Monitor token consumption, latency, API errors and cost in real time.

Usage tracking

Control API keys, tokens and AI costs in real time.

The usage management page lets engineering and finance teams track consumption in detail, including tokens used, resource reports and service billing and credits.

API keys per application
Token usage by plan
Usage reporting over time
Manage service payment/credit
FAQ

Frequently asked questions about HiTechCloud vMaaS.

What is vMaaS?

vMaaS delivers AI models as APIs, covering language, image, audio and embedding workloads, with OpenAI API-compatible endpoints.

Does the business need to invest in GPU/NPU hardware or an MLOps team?

No. You can integrate AI into applications, products or processes through the API without investing in AI infrastructure or building a dedicated MLOps team.

Is the service compatible with the OpenAI API?

Yes. vMaaS supports an OpenAI-compatible API standard, so engineering teams can integrate quickly with existing applications, SDKs or systems.

Which use cases is vMaaS suited to?

The service suits customer-support chatbots, knowledge base Q&A, document summarization, coding assistants, marketing content generation, RAG and AI-assisted process automation.

Is there a token-per-minute or API-request-per-minute limit?

vMaaS plans have no per-minute token or API request limits and do not price input and output tokens differently. Content length and maximum output are capped at 8K.

Is customer data used to train third-party models?

No. Customer data is not used to train third-party models, and the system is designed to run on HiTechCloud infrastructure in Vietnam.

Is BYOM supported?

Yes. The Customize plan includes advice on bespoke configurations and guidance on BYOM (Bring Your Own Model) deployment for organizations with specific requirements.

What payment methods are available?

The minimum term is one month. Payment is made in advance. Prices exclude VAT where applicable.

Ready to deploy vMaaS?

Integrate AI models into your applications with HiTechCloud.

HiTechCloud helps with model selection, API design, cost control and deploying AI at production scale.