- Add GeminiCloudModel: proxies inference to Gemini 2.0 Flash via HTTPS
using HttpURLConnection (no new deps). Supports both blocking and SSE
streaming. Works from any context — background, emulator, CI.
- Fix background inference: OnDeviceModel.create() now accepts an apiKey;
when set, GeminiCloudModel is selected immediately, bypassing Gemini
Nano's foreground-only restriction. ApiServerService reads the key from
SharedPreferences at startup.
- Add API key UI: password field in MainActivity saved to SharedPreferences
before the service starts; disabled while server is running.
- Add GitHub Actions CI (.github/workflows/ci.yml): build job produces a
debug APK artifact; test job spins up a KVM-accelerated Android 31
emulator, installs the APK, writes the GEMINI_API_KEY secret into
SharedPreferences via adb run-as, starts the service, and runs curl
assertions against /health, /v1/models, /v1/chat/completions, and the
SSE streaming endpoint.
- Add release workflow (.github/workflows/release.yml): triggered on v*
tags, builds the APK and creates a GitHub Release with it attached.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Android app that turns a Pixel 10 into a free AI API server by leveraging
the Tensor G5's on-device AI capabilities through an embedded HTTP server.
Key components:
- Dual AI backend: Gemini Nano (ML Kit Prompt API) + MediaPipe LLM fallback
- OpenAI-compatible REST API (chat/completions, completions, models, health)
- NanoHTTPD embedded server with CORS support and SSE streaming
- Foreground service with wake lock for persistent background operation
- Material Design 3 dashboard with live request logging
All inference runs entirely on-device — zero cloud costs, full offline capability.
https://claude.ai/code/session_01GvqMLSMmfMR8uz66BFVXX2