Skip to main content

Overview

@zerogpu/worker turns the Macs running your desktop app into compute nodes on the ZeroGPU network. It runs in its own background process (an Electron utility process), registers the Mac, downloads its assigned model, and serves inference tasks locally. There are no servers for you to run. Unlike the browser and Android SDKs, the worker doesn’t watch battery, sleep, or window focus. It keeps serving until you stop it, so your app decides when it runs. The recommended policy below runs it while your app is open and the Mac is on power.

Requirements

Installation

This adds the worker and its native inference runtimes (ONNX Runtime, and llama.cpp with Metal on Apple Silicon for GGUF models). The worker must ship as a real dependency in your packaged app’s node_modules; don’t bundle it into your main-process code.

Initialization

Create zerogpu-worker.mjs next to your main-process code. It starts the worker in a utility process with a free loopback port and a per-launch token, keeps its state in your app’s data directory, and stops it cleanly:
zerogpu-worker.mjs
The worker is configured entirely through environment variables:

Authentication

The worker authenticates with your SDK key (zgpu-sdk-…) through ZG_OPERATOR_KEY. The key only joins devices to your fleet, but it ships inside every copy of your app, so keep it out of source control and inject it at build time.
Never ship a ZeroGPU API key (zgpu-api-…) inside a desktop app. It authorizes paid inference for your whole organization, and anything in an app bundle can be extracted.

Register device

Registration is automatic. On start, the worker runs a short CPU benchmark (about 5 seconds), creates a random device ID stored in ZG_DATA_DIR, registers the Mac with ZeroGPU, downloads and verifies its assigned model, loads it, and opens an outbound connection for tasks. Expect under 30 seconds on first launch and a few seconds afterwards. All connections to ZeroGPU are outbound (HTTPS and WSS). The only listener is the local API on 127.0.0.1.

Start contributing compute

Create one ZeroGpuWorker in your main process and start it when the Mac is on power:
main.mjs

Stop / pause

worker.stop() sends SIGTERM: the worker leaves the network, flushes telemetry, unloads its model, and exits within 5 seconds. To pause, call stop(); to resume, call start(). The power policy above does exactly this on battery and sleep. Let the worker leave the network before your app quits:
main.mjs

Device lifecycle

Check status from your main process with await worker.status():
ready turns true once a model is loaded and has passed a warm-up inference. network.completedTasks counts tasks ZeroGPU sent to this Mac. Use the onExit option to show status or restart after a delay; the worker protects itself against crash loops with a cooldown.

Example integration

  1. Start your app with logLevel: "info".
  2. Open the worker log at ~/Library/Logs/<Your App>/zerogpu-worker.log. A healthy first start looks like this:
  3. Quit your app. The log ends with [lifecycle] Shutdown (host_stop).
Switch logLevel back to warn for release builds.

Packaging and signing

Native libraries must live outside the ASAR archive and be signed with your Developer ID. With electron-builder:
package.json
Build arm64 and x64 separately, each after npm ci --os=darwin --cpu=<arch>, so every architecture gets its matching native binaries.

Other app shells

Not using Electron? Any app that can ship a Node.js 20+ runtime can run the same entry point as a child process with the same variables:
Stop it with SIGTERM, allow 5 seconds for it to exit, and make sure it never outlives your app.

Troubleshooting