Skip to main content

Overview

The ZeroGPU Android SDK (ai.zerogpu:android-sdk) turns the devices running your app into on-device inference nodes on the ZeroGPU network. Once started, it registers the device, downloads the model the network assigns, and serves inference tasks locally with ONNX Runtime (and a bundled llama.cpp for GGUF models). The SDK follows your app’s lifecycle: it contributes compute while the app is in the foreground and disconnects when the app goes to the background.

Requirements

The SDK’s manifest declares INTERNET and ACCESS_NETWORK_STATE and sets android:largeHeap="true". These merge into your app automatically.

Installation

Make sure Maven Central is a repository:
settings.gradle.kts
Add the dependency to your app module:
build.gradle.kts
The SDK bundles native libraries (ONNX Runtime and llama.cpp). Add this to the android { } block so duplicate copies of libc++_shared.so don’t fail the build:
build.gradle.kts
Using Expo / React Native? Install the wrapper instead: npx expo install @zerogpu/expo-android-sdk, add "@zerogpu/expo-android-sdk" to your app config plugins, and run npx expo prebuild.

Initialization

Create one ZeroGpuSdk for your whole app, typically in Application.onCreate(). The base URL and environment default to production, so you only pass your SDK key:
Optional constructor parameters:

Authentication

The SDK authenticates with your SDK key (zgpu-sdk-…), passed as operatorKey. Don’t hard-code it. Put it in local.properties (git-ignored) and expose it through BuildConfig:
build.gradle.kts
local.properties

Register device

Registration is automatic. After start(), while the app is in the foreground, the SDK:
  1. Registers the device with the ZeroGPU device registry using your SDK key.
  2. Downloads its assigned model(s), verifies them, and caches them on device. Models are only re-downloaded when they change.
  3. Runs a warm-up inference.
  4. Opens a WebSocket and advertises the device as ready.

Start contributing compute

Call start() once. From then on the device manages itself: it follows the app in and out of the foreground, reconnects after network drops, and serves whatever tasks the network sends. You don’t poll it or feed it tasks.

Stop / pause

stop() tears everything down: it closes the connection, unloads models, and cancels in-flight work. There is no separate pause call; to pause, call stop(), and call start() again to resume. Going to the background already pauses contribution: the SDK closes its connection and reconnects when the app returns to the foreground.

Device lifecycle

Before downloading models, the SDK checks free memory and storage, and blocks downloads when the battery is below 20% and the device isn’t charging. Subscribe to lifecycle changes with a ZeroGpuListener. Callbacks arrive on a background thread:

Example integration

MainActivity.kt
To confirm the device is online, construct the SDK with enableLogs = true, run the app in the foreground, and watch logcat:
Once you see status":"idle" and WebSocket ready for tasks, the device is serving the network.

Troubleshooting