Skip to main content
The Model Catalog lets you browse every available ZeroGPU model and compare pricing across the tasks you care about. It’s useful when you’re selecting which model identifier to send to POST /v1/responses. Sections on this page: At a glance (pricing table), Detailed model cards, and Model library by task.

At a glance

Detailed model cards

deepseek-v4.1-flash
deepseek-v4.1-flash
1,048,576 context window$0.15 / 1M input$0.60 / 1M output$0.0025 / 1M cached input
DeepSeek’s DeepSeek-V4.1-Flash is an open-weight sparse Mixture-of-Experts model and the first built on DeepSeek’s Causal Encoder-Decoder (CED) architecture, activating 8B parameters on input and 16B on output. It…
glm-5.3-flash
glm-5.3-flash
1,048,576 context window$0.10 / 1M input$0.35 / 1M output$0.02 / 1M cached input
Z.ai’s GLM-5.3-Flash is an efficient open-weight model for coding and long-horizon agent tasks, served on ZeroGPU for general text generation. Its hybrid sparse and linear attention keeps…
gpt-5.4-nano
gpt-5.4-nano
400,000 context window$0.20 / 1M input$1.25 / 1M output$0.02 / 1M cached input
OpenAI’s GPT-5.4 nano is the most cost-efficient model in the GPT-5.4 family, served on ZeroGPU for high-volume and latency-sensitive workloads such as classification, extraction, routing, and…
gpt-5.6-luna
gpt-5.6-luna
272,000 context window$0.20 / 1M input$1.20 / 1M output$0.20 / 1M cached input
OpenAI’s GPT-5.6 Luna is the cost-optimized model of the GPT-5.6 family, served on ZeroGPU for cost-sensitive, high-volume workloads. It supports adjustable reasoning effort, function calling, and…
gpt-4.1-mini
gpt-4.1-mini
1,047,576 context window$0.40 / 1M input$1.60 / 1M output$0.10 / 1M cached input
OpenAI’s GPT-4.1 mini is the fast, cost-efficient model of the GPT-4.1 family, served on ZeroGPU for general text generation. It excels at instruction following and tool calling, supports function…
deepseek-v4-flash-0731
deepseek-v4-flash-0731
1,048,576 context window$0.16 / 1M input$0.38 / 1M output$0.006 / 1M cached input
DeepSeek’s DeepSeek-V4-Flash is an open-weight Mixture-of-Experts model built for efficient reasoning, coding, and agentic workflows, with 284B total parameters activating only 13B per token. Its…
gpt-oss-120b
gpt-oss-120b
131,072 context window$0.15 / 1M input$0.60 / 1M output$0.03 / 1M cached input
OpenAI’s gpt-oss-120b is an open-weight Mixture-of-Experts model with 117B total parameters (5.1B active per token), served on ZeroGPU for general text generation. It reasons through a problem…
qwen3-30b-a3b-fp8
qwen3-30b-a3b-fp8
32,768 context window$0.10 / 1M input$0.45 / 1M output$0.05 / 1M cached input
Alibaba’s Qwen3-30B-A3B is an open-weight Mixture-of-Experts model with 30.5B total parameters (3.3B active per token), served on ZeroGPU as an FP8 build for efficient inference. It thinks through a problem…
glm-5.2
glm-5.2
262,144 context window$1.10 / 1M input$3.50 / 1M output$0.26 / 1M cached input
Z.ai’s GLM-5.2 is an open-weight Mixture-of-Experts flagship built for long-horizon tasks, with 753B total parameters activating 8 of 256 experts per token. It sustains a 262,144-token (256K)…
llama-3.1-8b-instruct-fast
llama-3.1-8b-instruct-fast
131,072 max tokens$0.15 / 1M input$0.28 / 1M output$0.025 / 1M cached input
Meta’s Llama 3.1 Instruct, tuned for fast, low-cost summarization at scale on the ZeroGPU edge network. Its 128K-token context window takes in entire documents, long transcripts, and full email or…
llama-guard-4-12b
llama-guard-4-12b
Text GenerationSafety classificationText ModerationBrand Safety12B params163,840 context window$0.18 / 1M input$0.18 / 1M output
Meta’s Llama Guard 4 12B is a multimodal safety classification model for moderating text, images, and mixed text-image inputs. It evaluates both incoming prompts…
zlm-v2-iab-classify-edge-enriched
zlm-v2-iab-classify-edge-enriched
800 max tokens$0.025 / 1M input$0.15 / 1M output
The enriched variant of ZeroGPU’s IAB classifier turns a single inference call into a full content-intelligence profile not just a label, but everything a contextual pipeline needs to act on. Each…
zlm-v1-iab-classify-edge
zlm-v1-iab-classify-edge
400 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s IAB classifier maps any text to the industry-standard IAB Content Taxonomy in a single, fast inference call. Each call returns categories across both the 1.0 and 2.2 taxonomies plus matched…
zlm-v1-iab-domain-classifier
zlm-v1-iab-domain-classifier
100 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s Domain IAB Classifier maps a raw domain name straight to the IAB Content Taxonomy, returning content categories, topics, keywords, and user-intent signals. It needs only the domain as input, cutting payload size by up to 10x versus page-level…
zlm-v1-moderation-edge
zlm-v1-moderation-edge
800 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s moderation model returns the complete OpenAI 13-category taxonomy — a flagged verdict, category booleans, and calibrated scores — as a drop-in for omni-moderation-latest. In head-to-head benchmarks it wins the binary safe/unsafe decision (0.899 vs 0.853 F1) and 9 of 13 harm categories, and returns verdicts 1.2–1.8× faster on production-range…
gliner-multi-pii-v1
gliner-multi-pii-v1
800 max tokens$0.02 / 1M input$0.05 / 1M output
GLiNER Multi PII is a multilingual PII detection and redaction model that supports on-prem deployments as well. It identifies 40+ personally identifiable entity types — identity, contact, government…
gliner2-base-v1
gliner2-base-v1
800 max tokens$0.02 / 1M input$0.05 / 1M output
gliner2-base-v1 is a versatile extraction-and-classification model for the structured tasks that fill most production pipelines. Point it at any text and, with a single API call, pull named entities…
deberta-v3-small
deberta-v3-small
400 max tokens$0.02 / 1M input$0.05 / 1M output
Microsoft’s DeBERTa-v3-small is a fast, lightweight zero-shot text classifier for high-volume routing, filtering, and tagging. Hand it any text alongside your own candidate labels and it returns a…
LFM2.5-1.2B-Thinking
LFM2.5-1.2B-Thinking
32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Thinking is a compact reasoning model that works through a problem step by step before it answers. Built by Liquid AI, it generates an explicit chain-of-thought trace so for…
LFM2.5-1.2B-Instruct
LFM2.5-1.2B-Instruct
32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Instruct is a hybrid architecture model purpose-built for on-device deployment, trained on 28 trillion tokens with multi-stage reinforcement learning. It delivers…
zlm-v1-signal-extract
zlm-v1-signal-extract
Text ClassificationSignal extractionAd TechChat80M params400 context window$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s signal extractor turns unstructured text into structured signals that downstream systems can act on. One inference call returns topics, keywords, intent…
all-minilm-l6-v2
all-minilm-l6-v2
512 max tokens$0.004 / 1M input
Sentence-Transformers’ all-MiniLM-L6-v2 is the default workhorse of semantic search. It maps a sentence or short paragraph to a 384-dimensional vector, trained with contrastive learning on more than a billion sentence pairs…
bge-small-en-v1.5
bge-small-en-v1.5
512 max tokens$0.004 / 1M input
BAAI’s BGE-small-en-v1.5 is a retrieval-first English embedding model and one of the strongest performers on the MTEB benchmark for its size. It produces the same 384-dimensional vectors, with a 512-token window…

Model library by task

Open Weight

8 models: deepseek-v4.1-flash, glm-5.3-flash, deepseek-v4-flash-0731, glm-5.2, qwen3-30b-a3b-fp8, gpt-oss-120b, llama-3.1-8b-instruct-fast, llama-guard-4-12b

Text Generation

13 models: deepseek-v4.1-flash, glm-5.3-flash, gpt-5.4-nano, gpt-5.6-luna, gpt-4.1-mini, deepseek-v4-flash-0731, gpt-oss-120b, qwen3-30b-a3b-fp8, glm-5.2, llama-3.1-8b-instruct-fast, llama-guard-4-12b, LFM2.5-1.2B-Thinking, LFM2.5-1.2B-Instruct

Text Classification

5 models: zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-classify-edge, zlm-v1-iab-domain-classifier, deberta-v3-small, zlm-v1-signal-extract

Moderation

1 model: zlm-v1-moderation-edge

Text Embedding

2 models: all-minilm-l6-v2, bge-small-en-v1.5

Data Extraction

2 models: gliner-multi-pii-v1, gliner2-base-v1

Ad Tech

4 models: zlm-v1-iab-classify-edge, zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-domain-classifier, zlm-v1-signal-extract