Skip to main content
Moderation models screen text for unsafe, harmful, or policy-sensitive content before it reaches your application or model workflow. Call either the Responses API (POST /v1/responses) or the OpenAI-compatible Chat Completions API (POST /v1/chat/completions).

zlm-v1-moderation-edge

ZeroGPU’s moderation model screens text for unsafe, harmful, or policy-sensitive content and returns the complete OpenAI 13-category taxonomy — a flagged verdict, per-category booleans, and calibrated category_scores — so it drops into any pipeline written against omni-moderation-latest. Under the hood it’s an 86M-parameter DeBERTa encoder with a shared trunk feeding one binary safe/unsafe head and 13 category heads, with per-category thresholds calibrated on held-out validation data. In head-to-head benchmarks against OpenAI omni-moderation it wins the binary safe/unsafe decision (0.899 vs 0.853 F1) and 9 of 13 harm categories — with the largest gains on graphic violence, illicit content, and self-harm — while returning verdicts 1.2–1.8× faster at the median on production-range inputs, because inference is co-located at the edge instead of a round trip to a central API. Moderation sits inline in front of every response your app serves; this is the model that’s fast and accurate enough to live there.
References: Moderation benchmarkTermsPrivacy
The API returns a standard completion envelope. The moderation verdict is inside it as an escaped JSON string — not a nested object — in choices[].message.content (Chat Completions) or output[].content[].text (Responses API). This is the actual, unmodified output:
JSON-parse that string to get the moderation result — flagged, unsafe_score, and the 13 categories with their category_scores:
Parsed moderation result