model identifier to send to POST /v1/responses.
Sections on this page: At a glance (pricing table), Detailed model cards, and Model library by task.
At a glance
Detailed model cards
deepseek-v4.1-flash
1,048,576 context window$0.15 / 1M input$0.60 / 1M output$0.0025 / 1M cached input
DeepSeek’s DeepSeek-V4.1-Flash is an open-weight sparse Mixture-of-Experts model and the first built on DeepSeek’s Causal Encoder-Decoder (CED) architecture, activating 8B parameters on input and 16B on output. It…glm-5.3-flash
1,048,576 context window$0.10 / 1M input$0.35 / 1M output$0.02 / 1M cached input
Z.ai’s GLM-5.3-Flash is an efficient open-weight model for coding and long-horizon agent tasks, served on ZeroGPU for general text generation. Its hybrid sparse and linear attention keeps…gpt-5.4-nano
400,000 context window$0.20 / 1M input$1.25 / 1M output$0.02 / 1M cached input
OpenAI’s GPT-5.4 nano is the most cost-efficient model in the GPT-5.4 family, served on ZeroGPU for high-volume and latency-sensitive workloads such as classification, extraction, routing, and…gpt-5.6-luna
272,000 context window$0.20 / 1M input$1.20 / 1M output$0.20 / 1M cached input
OpenAI’s GPT-5.6 Luna is the cost-optimized model of the GPT-5.6 family, served on ZeroGPU for cost-sensitive, high-volume workloads. It supports adjustable reasoning effort, function calling, and…gpt-4.1-mini
1,047,576 context window$0.40 / 1M input$1.60 / 1M output$0.10 / 1M cached input
OpenAI’s GPT-4.1 mini is the fast, cost-efficient model of the GPT-4.1 family, served on ZeroGPU for general text generation. It excels at instruction following and tool calling, supports function…deepseek-v4-flash-0731
1,048,576 context window$0.16 / 1M input$0.38 / 1M output$0.006 / 1M cached input
DeepSeek’s DeepSeek-V4-Flash is an open-weight Mixture-of-Experts model built for efficient reasoning, coding, and agentic workflows, with 284B total parameters activating only 13B per token. Its…gpt-oss-120b
131,072 context window$0.15 / 1M input$0.60 / 1M output$0.03 / 1M cached input
OpenAI’s gpt-oss-120b is an open-weight Mixture-of-Experts model with 117B total parameters (5.1B active per token), served on ZeroGPU for general text generation. It reasons through a problem…qwen3-30b-a3b-fp8
32,768 context window$0.10 / 1M input$0.45 / 1M output$0.05 / 1M cached input
Alibaba’s Qwen3-30B-A3B is an open-weight Mixture-of-Experts model with 30.5B total parameters (3.3B active per token), served on ZeroGPU as an FP8 build for efficient inference. It thinks through a problem…glm-5.2
262,144 context window$1.10 / 1M input$3.50 / 1M output$0.26 / 1M cached input
Z.ai’s GLM-5.2 is an open-weight Mixture-of-Experts flagship built for long-horizon tasks, with 753B total parameters activating 8 of 256 experts per token. It sustains a 262,144-token (256K)…llama-3.1-8b-instruct-fast
131,072 max tokens$0.15 / 1M input$0.28 / 1M output$0.025 / 1M cached input
Meta’s Llama 3.1 Instruct, tuned for fast, low-cost summarization at scale on the ZeroGPU edge network. Its 128K-token context window takes in entire documents, long transcripts, and full email or…llama-guard-4-12b
Text GenerationSafety classificationText ModerationBrand Safety12B params163,840 context window$0.18 / 1M input$0.18 / 1M output
Meta’s Llama Guard 4 12B is a multimodal safety classification model for moderating text, images, and mixed text-image inputs. It evaluates both incoming prompts…zlm-v2-iab-classify-edge-enriched
800 max tokens$0.025 / 1M input$0.15 / 1M output
The enriched variant of ZeroGPU’s IAB classifier turns a single inference call into a full content-intelligence profile not just a label, but everything a contextual pipeline needs to act on. Each…zlm-v1-iab-classify-edge
400 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s IAB classifier maps any text to the industry-standard IAB Content Taxonomy in a single, fast inference call. Each call returns categories across both the 1.0 and 2.2 taxonomies plus matched…zlm-v1-iab-domain-classifier
100 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s Domain IAB Classifier maps a raw domain name straight to the IAB Content Taxonomy, returning content categories, topics, keywords, and user-intent signals. It needs only the domain as input, cutting payload size by up to 10x versus page-level…zlm-v1-moderation-edge
800 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s moderation model returns the complete OpenAI 13-category taxonomy — a flagged verdict, category booleans, and calibrated scores — as a drop-in for omni-moderation-latest. In head-to-head benchmarks it wins the binary safe/unsafe decision (0.899 vs 0.853 F1) and 9 of 13 harm categories, and returns verdicts 1.2–1.8× faster on production-range…gliner-multi-pii-v1
800 max tokens$0.02 / 1M input$0.05 / 1M output
GLiNER Multi PII is a multilingual PII detection and redaction model that supports on-prem deployments as well. It identifies 40+ personally identifiable entity types — identity, contact, government…gliner2-base-v1
800 max tokens$0.02 / 1M input$0.05 / 1M output
gliner2-base-v1 is a versatile extraction-and-classification model for the structured tasks that fill most production pipelines. Point it at any text and, with a single API call, pull named entities…deberta-v3-small
400 max tokens$0.02 / 1M input$0.05 / 1M output
Microsoft’s DeBERTa-v3-small is a fast, lightweight zero-shot text classifier for high-volume routing, filtering, and tagging. Hand it any text alongside your own candidate labels and it returns a…LFM2.5-1.2B-Thinking
32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Thinking is a compact reasoning model that works through a problem step by step before it answers. Built by Liquid AI, it generates an explicit chain-of-thought trace so for…LFM2.5-1.2B-Instruct
32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Instruct is a hybrid architecture model purpose-built for on-device deployment, trained on 28 trillion tokens with multi-stage reinforcement learning. It delivers…zlm-v1-signal-extract
Text ClassificationSignal extractionAd TechChat80M params400 context window$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s signal extractor turns unstructured text into structured signals that downstream systems can act on. One inference call returns topics, keywords, intent…all-minilm-l6-v2
512 max tokens$0.004 / 1M input
Sentence-Transformers’ all-MiniLM-L6-v2 is the default workhorse of semantic search. It maps a sentence or short paragraph to a 384-dimensional vector, trained with contrastive learning on more than a billion sentence pairs…bge-small-en-v1.5
512 max tokens$0.004 / 1M input
BAAI’s BGE-small-en-v1.5 is a retrieval-first English embedding model and one of the strongest performers on the MTEB benchmark for its size. It produces the same 384-dimensional vectors, with a 512-token window…Model library by task
Open Weight
8 models:
deepseek-v4.1-flash, glm-5.3-flash, deepseek-v4-flash-0731, glm-5.2, qwen3-30b-a3b-fp8, gpt-oss-120b, llama-3.1-8b-instruct-fast, llama-guard-4-12bText Generation
13 models:
deepseek-v4.1-flash, glm-5.3-flash, gpt-5.4-nano, gpt-5.6-luna, gpt-4.1-mini, deepseek-v4-flash-0731, gpt-oss-120b, qwen3-30b-a3b-fp8, glm-5.2, llama-3.1-8b-instruct-fast, llama-guard-4-12b, LFM2.5-1.2B-Thinking, LFM2.5-1.2B-InstructText Classification
5 models:
zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-classify-edge, zlm-v1-iab-domain-classifier, deberta-v3-small, zlm-v1-signal-extractModeration
1 model:
zlm-v1-moderation-edgeText Embedding
2 models:
all-minilm-l6-v2, bge-small-en-v1.5Data Extraction
2 models:
gliner-multi-pii-v1, gliner2-base-v1Ad Tech
4 models:
zlm-v1-iab-classify-edge, zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-domain-classifier, zlm-v1-signal-extract
