Skip to main content
The Model Catalog lets you browse every available ZeroGPU model and compare pricing across the tasks you care about. It’s useful when you’re selecting which model identifier to send to POST /v1/responses. Sections on this page: At a glance (pricing table), Detailed model cards, and Model library by task.

At a glance

Detailed model cards

gpt-oss-120b
gpt-oss-120b
Text GenerationReasoningFunction callingBatch tasks117B params131,072 context window$0.03 / 1M input$0.10 / 1M output
OpenAI’s gpt-oss-120b is an open-weight Mixture-of-Experts model with 117B total parameters (5.1B active per token), served on ZeroGPU for general text generation. It reasons through a problem…
qwen3-30b-a3b-fp8
qwen3-30b-a3b-fp8
Text GenerationReasoningFunction callingStreamingBatch tasksMultilingual30.5B params32,768 context window$0.05 / 1M input$0.30 / 1M output
Alibaba’s Qwen3-30B-A3B is an open-weight Mixture-of-Experts model with 30.5B total parameters (3.3B active per token), served on ZeroGPU as an FP8 build for efficient inference. It thinks through a problem…
llama-3.1-8b-instruct-fast
llama-3.1-8b-instruct-fast
Summarization8B params131,072 max tokens$0.02 / 1M input$0.05 / 1M output
Meta’s Llama 3.1 Instruct, tuned for fast, low-cost summarization at scale on the ZeroGPU edge network. Its 128K-token context window takes in entire documents, long transcripts, and full email or…
zlm-v2-iab-classify-edge-enriched
zlm-v2-iab-classify-edge-enriched
Text Classification90M params400 max tokens$0.02 / 1M input$0.05 / 1M output
The enriched variant of ZeroGPU’s IAB classifier turns a single inference call into a full content-intelligence profile not just a label, but everything a contextual pipeline needs to act on. Each…
zlm-v1-iab-classify-edge
zlm-v1-iab-classify-edge
Text Classification90M params400 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s IAB classifier maps any text to the industry-standard IAB Content Taxonomy in a single, fast inference call. Each call returns categories across both the 1.0 and 2.2 taxonomies plus matched…
zlm-v1-iab-domain-classifier
zlm-v1-iab-domain-classifier
Text Classification149M params100 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s Domain IAB Classifier maps a raw domain name straight to the IAB Content Taxonomy, returning content categories, topics, keywords, and user-intent signals. It needs only the domain as input, cutting payload size by up to 10x versus page-level…
zlm-v1-followup-questions-edge
zlm-v1-followup-questions-edge
Text Generation120M params400 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s follow-up question generator takes any piece of content — an article, an answer, a chat turn and returns a short set of natural questions a reader would actually ask next. Where a…
gliner-multi-pii-v1
gliner-multi-pii-v1
PII300M params800 max tokens$0.02 / 1M input$0.05 / 1M output
GLiNER Multi PII is a multilingual PII detection and redaction model that supports on-prem deployments as well. It identifies 40+ personally identifiable entity types — identity, contact, government…
gliner2-base-v1
gliner2-base-v1
PII205M params800 max tokens$0.02 / 1M input$0.05 / 1M output
gliner2-base-v1 is a versatile extraction-and-classification model for the structured tasks that fill most production pipelines. Point it at any text and, with a single API call, pull named entities…
deberta-v3-small
deberta-v3-small
Text Classification142M params400 max tokens$0.02 / 1M input$0.05 / 1M output
Microsoft’s DeBERTa-v3-small is a fast, lightweight zero-shot text classifier for high-volume routing, filtering, and tagging. Hand it any text alongside your own candidate labels and it returns a…
LFM2.5-1.2B-Thinking
LFM2.5-1.2B-Thinking
Text Generation1.2B params32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Thinking is a compact reasoning model that works through a problem step by step before it answers. Built by Liquid AI, it generates an explicit chain-of-thought trace so for…
LFM2.5-1.2B-Instruct
LFM2.5-1.2B-Instruct
Text Generation1.2B params32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Instruct is a hybrid architecture model purpose-built for on-device deployment, trained on 28 trillion tokens with multi-stage reinforcement learning. It delivers…

Model library by task

Ad Tech

3 models: zlm-v1-iab-classify-edge, zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-domain-classifier

Text Classification

4 models: zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-classify-edge, zlm-v1-iab-domain-classifier, deberta-v3-small

Text Generation

5 models: gpt-oss-120b, qwen3-30b-a3b-fp8, zlm-v1-followup-questions-edge, LFM2.5-1.2B-Thinking, LFM2.5-1.2B-Instruct

PII

2 models: gliner-multi-pii-v1, gliner2-base-v1

Summarization

1 model: llama-3.1-8b-instruct-fast