Files
openclaw/docs/testing/epic-305-voice-call-test-plan.md
Clawd BotandClaude Opus 4.6 ca9b510922 chore: align with upstream openclaw/openclaw and overlay local additions
- Reset master to upstream/main (16,697 commits)
- Overlay 2,271 local-only files (skills, tools, workspace, configs, apps)
- Restore IDENTITY.md and USER.md templates
- Build verified, gateway running, Discord working

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-03-03 07:40:46 +01:00

20 KiB

title, summary
title summary
Epic 305: Voice Call Testing Plan Manual test plan, automated smoke tests, and HTTPS setup for voice call functionality

Epic 305: Voice Call Testing Plan

Overview

This document covers manual testing procedures, expected latencies, failure modes, and automated test requirements for the voice-call plugin, including browser-based calling with getUserMedia.

Prerequisites

HTTPS Requirements for getUserMedia

Modern browsers require HTTPS for getUserMedia() API access, with these exceptions:

  • localhost (always permitted)
  • 127.0.0.1 (always permitted)
  • LAN access from non-localhost IPs requires HTTPS

mkcert Setup for Local Development

When testing on a LAN from non-localhost addresses (e.g., accessing your dev machine from a phone on the same network):

Install mkcert

macOS:

brew install mkcert
mkcert -install

Linux:

# Install certutil (required for mkcert)
sudo apt install libnss3-tools  # Debian/Ubuntu
# or
sudo yum install nss-tools      # Fedora/CentOS

# Install mkcert
curl -JLO "https://dl.filippo.io/mkcert/latest?for=linux/amd64"
chmod +x mkcert-v*-linux-amd64
sudo mv mkcert-v*-linux-amd64 /usr/local/bin/mkcert
mkcert -install

Windows:

choco install mkcert
mkcert -install

Generate Local Certificates

# Navigate to your project
cd /path/to/openclaw

# Generate cert for your local IP (find with `ip addr` or `ifconfig`)
mkcert localhost 127.0.0.1 192.168.1.100 ::1

# This creates:
# - localhost+3.pem (certificate)
# - localhost+3-key.pem (private key)

Configure Gateway to Use HTTPS

Update your OpenClaw config:

{
  gateway: {
    server: {
      https: {
        enabled: true,
        cert: "/path/to/localhost+3.pem",
        key: "/path/to/localhost+3-key.pem",
      },
    },
  },
}

Trust Certificate on Test Devices

iOS/Android (testing from phone):

  1. Email yourself the root CA: cat "$(mkcert -CAROOT)/rootCA.pem"
  2. Install the certificate on your device
  3. iOS: Settings → General → About → Certificate Trust Settings → Enable
  4. Android: Settings → Security → Install from storage

Other computers on LAN:

# Copy the CA cert
scp "$(mkcert -CAROOT)/rootCA.pem" user@test-machine:~/
# On test machine, install it
mkcert -install ~/rootCA.pem

Manual Test Plan

Test Environment Setup

Component Requirement
Gateway Running locally or on accessible host
Network Public webhook URL (ngrok/Tailscale Funnel) or stable HTTPS endpoint
Provider Twilio, Telnyx, or Plivo account with credits
Phone Test phone number for inbound/outbound calls
Browser Chrome/Firefox/Safari with HTTPS enabled

Test Cases

TC-001: Plugin Installation & Startup

Objective: Verify plugin loads and initializes correctly

Steps:

  1. Install plugin: openclaw plugins install @openclaw/voice-call
  2. Configure provider credentials (see config section below)
  3. Restart gateway: openclaw gateway restart
  4. Verify plugin loaded: openclaw channels status --probe

Expected Result:

  • Plugin appears in status output
  • No errors in gateway logs
  • Webhook server starts on configured port

Expected Latency: < 5 seconds for gateway restart

Failure Modes:

  • Missing credentials → Error on startup with clear message
  • Port already in use → Bind error with port number
  • Invalid provider config → Validation error with field name

TC-002: Webhook Exposure

Objective: Verify webhook endpoint is publicly reachable

Steps:

  1. Expose via Tailscale: openclaw voicecall expose --mode funnel
    • OR configure ngrok tunnel in config
    • OR set static publicUrl in config
  2. Verify endpoint: curl -X POST https://your-public-url/voice/webhook
  3. Check signature verification doesn't crash on invalid payload

Expected Result:

  • Webhook responds with 400/405 for invalid requests
  • Signature verification rejects unsigned requests (when not in bypass mode)
  • Logs show webhook received

Expected Latency: < 200ms for webhook response

Failure Modes:

  • Tunnel fails to establish → Clear error message with provider name
  • Certificate issues → TLS error, suggest mkcert or valid cert
  • Firewall blocks → Timeout, suggest checking firewall rules

TC-003: Outbound Call - Notify Mode

Objective: Place an outbound notification call

Steps:

  1. Configure provider and fromNumber in config
  2. Execute: openclaw voicecall call --to "+15555550123" --message "Test notification" --mode notify
  3. Answer the call on receiving phone
  4. Listen to the message
  5. Hang up

Expected Result:

  • Call initiated within 3 seconds
  • callId returned immediately
  • TTS plays configured message
  • Call ends automatically after message completes
  • Status shows completed

Expected Latencies:

Stage Latency
API call to provider 500ms - 2s
Ring time 5s - 30s (network dependent)
TTS generation 1s - 5s (depending on provider/length)
Audio playback Real-time (message duration)

Failure Modes:

  • Invalid phone number → Provider error, status failed
  • Insufficient credits → Provider 402/403 error
  • TTS provider down → Fallback to provider native voice (if configured)
  • Webhook signature fails → Call proceeds but events not processed

TC-004: Outbound Call - Conversation Mode

Objective: Place an outbound call with multi-turn conversation

Steps:

  1. Execute: openclaw voicecall call --to "+15555550123" --message "Hello, can you hear me?" --mode conversation
  2. Answer the call
  3. Speak a response: "Yes, I can hear you"
  4. Wait for AI response
  5. Continue: openclaw voicecall continue --call-id <id> --message "Great, ending now"
  6. Verify call ends gracefully

Expected Result:

  • Initial message plays
  • User speech transcribed via STT
  • Agent responds appropriately
  • Follow-up messages play in order
  • Call ends on command

Expected Latencies:

Stage Latency
Speech detection start 200ms - 1s
Transcription (streaming) Real-time + 500ms
Agent inference 1s - 10s (model dependent)
TTS generation 1s - 5s
Audio playback start < 500ms

Failure Modes:

  • STT provider unreachable → Timeout, call falls back to notify mode
  • Agent model timeout → Default error message plays
  • Network interruption → Call drops, status failed
  • Concurrent TTS queue overflow → Messages queued, may delay

TC-005: Inbound Call Handling

Objective: Receive and process inbound calls

Steps:

  1. Configure inbound policy:
    {
      inboundPolicy: "allowlist",
      allowFrom: ["+15555550123"],
      inboundGreeting: "Hello! You've reached OpenClaw.",
    }
    
  2. Call the configured fromNumber from allowlisted number
  3. Wait for greeting
  4. Speak: "What's the weather?"
  5. Wait for response
  6. Hang up

Expected Result:

  • Call answered automatically
  • Greeting plays
  • User speech processed
  • Agent responds contextually
  • Call ends when user hangs up

Expected Latencies:

  • Same as TC-004 (conversation mode)

Failure Modes:

  • Call from non-allowlisted number → Rejected (400/403)
  • Policy set to disabled → Call not answered
  • Webhook down → Provider retries, eventual timeout

TC-006: WebSocket Media Streaming (Twilio)

Objective: Verify real-time media streaming for low-latency interaction

Prerequisites:

  • Twilio provider configured
  • streaming.enabled: true in config
  • OpenAI Realtime API key set

Steps:

  1. Initiate call with streaming enabled
  2. Answer call
  3. Speak immediately after greeting
  4. Verify partial transcripts appear in logs
  5. Check TTS interruption works (speak while agent is talking)

Expected Result:

  • WebSocket connection established
  • Audio chunks stream bidirectionally
  • Partial transcripts visible in near real-time
  • TTS queue clears on interruption

Expected Latencies:

Stage Latency
WebSocket handshake < 500ms
Audio chunk delivery 20ms - 100ms
Partial transcript 200ms - 1s
Interruption detection < 500ms

Failure Modes:

  • WebSocket upgrade fails → Falls back to HTTP webhooks + polling
  • OpenAI Realtime API down → STT timeout, call may fail
  • Network packet loss → Audio glitches, transcription errors

TC-007: Call Status & Monitoring

Objective: Verify call state tracking

Steps:

  1. Start a call: openclaw voicecall call --to "+15555550123" --message "Status test"
  2. Immediately check status: openclaw voicecall status --call-id <id>
  3. Tail events: openclaw voicecall tail (in separate terminal)
  4. Answer call, let message play, hang up
  5. Check final status

Expected Result:

  • Status shows progression: initiated → ringing → in-progress → completed
  • Tail shows real-time events
  • Final status includes duration, providerCallId

Expected Latency: Status queries < 100ms (local lookup)

Failure Modes:

  • Call state stuck → Event processing failed, check logs
  • Provider ID mismatch → Call not found by provider ID

TC-008: Error Recovery & Timeouts

Objective: Verify graceful handling of failures

Test Scenarios:

A. Provider API Timeout

  1. Simulate slow network (e.g., tc qdisc add dev eth0 root netem delay 5000ms)
  2. Initiate call
  3. Verify timeout error

Expected: Call marked failed within 30s, clear error message

B. Invalid Credentials

  1. Set incorrect authToken
  2. Attempt call
  3. Verify error

Expected: 401/403 from provider, call not initiated, credentials error logged

C. Webhook Signature Failure

  1. Set correct credentials but wrong publicUrl
  2. Initiate call
  3. Check logs

Expected: Call proceeds (provider doesn't know), but webhook events rejected, signature mismatch logged

D. TTS Provider Down

  1. Set invalid OpenAI API key
  2. Initiate call
  3. Verify fallback

Expected: Falls back to provider native voice (Twilio/Telnyx/Plivo TTS)


Browser-Based Calling (getUserMedia)

TC-009: Browser Voice Call (HTTPS Required)

Objective: Verify browser-based calling with getUserMedia

Prerequisites:

  • Gateway running on HTTPS (see mkcert setup above)
  • Browser with microphone permissions

Steps:

  1. Navigate to web UI: https://your-host:port/voice (or wherever UI is exposed)
  2. Click "Start Call" button
  3. Grant microphone permission when prompted
  4. Speak test phrase
  5. Verify audio transmission
  6. End call

Expected Result:

  • getUserMedia permission prompt appears
  • Microphone access granted
  • Audio visualizer shows levels
  • Speech transcribed correctly
  • Call ends cleanly

Expected Latencies:

  • getUserMedia permission: Instant (if previously granted)
  • Audio stream start: < 500ms
  • WebRTC connection: 1s - 3s
  • End-to-end latency: 300ms - 1s

Failure Modes:

  • HTTP instead of HTTPS → getUserMedia throws NotAllowedError
  • Microphone blocked → Permission denied error
  • Certificate untrusted → Browser warning, connection blocked
  • Network change → ICE reconnection or call drop

HTTPS Troubleshooting:

  • Verify URL starts with https://
  • Check browser console for mixed content warnings
  • Ensure certificate is trusted (see mkcert setup)
  • Test from localhost first (always works), then LAN IP

Automated Test Suite

Smoke Tests (Required for CI)

Location: extensions/voice-call/src/smoke.test.ts

import { describe, it, expect, beforeAll, afterAll } from "vitest";
import { VoiceCallManager } from "./manager.js";
import { MockProvider } from "./providers/mock.js";
import WebSocket from "ws";

describe("Voice Call - Smoke Tests", () => {
  let manager: VoiceCallManager;
  let webhookServer: any;

  beforeAll(async () => {
    // Initialize manager with mock provider
    manager = new VoiceCallManager({
      enabled: true,
      provider: "mock",
      fromNumber: "+15550000000",
      serve: { port: 0 }, // random port
    });

    webhookServer = await manager.startWebhookServer();
  });

  afterAll(async () => {
    await webhookServer?.close();
  });

  it("server starts and binds to port", () => {
    expect(webhookServer).toBeDefined();
    expect(webhookServer.listening).toBe(true);
    expect(webhookServer.address().port).toBeGreaterThan(0);
  });

  it("websocket upgrade succeeds", async () => {
    const port = webhookServer.address().port;
    const ws = new WebSocket(`ws://localhost:${port}/voice/stream`);

    await new Promise((resolve, reject) => {
      ws.on("open", resolve);
      ws.on("error", reject);
      setTimeout(() => reject(new Error("Timeout")), 5000);
    });

    expect(ws.readyState).toBe(WebSocket.OPEN);
    ws.close();
  });

  it("agent command invoked on call initiation", async () => {
    const { callId, success } = await manager.initiateCall("+15550001111", undefined, {
      message: "Test",
      mode: "notify",
    });

    expect(success).toBe(true);
    expect(callId).toMatch(/^call-/);

    const call = manager.getCall(callId);
    expect(call?.status).toBe("initiated");
  });

  it("call state transitions correctly", async () => {
    const { callId } = await manager.initiateCall("+15550001111");

    // Simulate provider events
    manager.processEvent({
      id: "evt-1",
      type: "call.ringing",
      callId,
      providerCallId: "provider-123",
      timestamp: Date.now(),
    });

    expect(manager.getCall(callId)?.status).toBe("ringing");

    manager.processEvent({
      id: "evt-2",
      type: "call.answered",
      callId,
      providerCallId: "provider-123",
      timestamp: Date.now(),
    });

    expect(manager.getCall(callId)?.status).toBe("in-progress");
  });

  it("webhook signature verification runs", () => {
    const provider = new MockProvider();
    const result = provider.verifyWebhook({
      method: "POST",
      url: "/voice/webhook",
      headers: {},
      body: {},
      rawBody: "",
    });

    expect(result).toHaveProperty("ok");
  });
});

Integration Tests (Optional, Requires Live Credentials)

Location: extensions/voice-call/src/integration.test.ts

import { describe, it, expect, beforeAll } from "vitest";

describe("Voice Call - Integration (Live)", () => {
  const LIVE = process.env.VOICE_CALL_LIVE_TEST === "1";

  it.skipIf(!LIVE)("places real outbound call", async () => {
    // Requires TWILIO_ACCOUNT_SID, TWILIO_AUTH_TOKEN, etc.
    // Test against real provider APIs
  });
});

Test Coverage Requirements

  • Minimum coverage: 70% (lines/branches/functions/statements)
  • Critical paths: 90%+ coverage
    • Call state machine
    • Event processing
    • Webhook verification
    • TTS queue management

Configuration Examples

Minimal Local Dev (Mock Provider)

{
  plugins: {
    entries: {
      "voice-call": {
        enabled: true,
        config: {
          provider: "mock",
          fromNumber: "+15550000000",
          serve: { port: 3334 },
        },
      },
    },
  },
}

Twilio with ngrok (Local Dev)

{
  plugins: {
    entries: {
      "voice-call": {
        enabled: true,
        config: {
          provider: "twilio",
          fromNumber: "+15550001234",
          twilio: {
            accountSid: "ACxxxxxxxx",
            authToken: "your_token",
          },
          serve: { port: 3334 },
          tunnel: {
            provider: "ngrok",
            allowNgrokFreeTierLoopbackBypass: true,
          },
          streaming: {
            enabled: true,
            streamPath: "/voice/stream",
          },
        },
      },
    },
  },
}

Production with Tailscale Funnel

{
  plugins: {
    entries: {
      "voice-call": {
        enabled: true,
        config: {
          provider: "twilio",
          fromNumber: "+15550001234",
          twilio: {
            accountSid: "ACxxxxxxxx",
            authToken: "your_token",
          },
          serve: {
            port: 3334,
            path: "/voice/webhook",
          },
          tailscale: {
            mode: "funnel",
            path: "/voice/webhook",
          },
          inboundPolicy: "allowlist",
          allowFrom: ["+15550005678"],
          streaming: { enabled: true },
        },
      },
    },
  },
}

Monitoring & Debugging

Log Locations

  • Gateway logs: ~/.openclaw/logs/gateway.log
  • Voice call plugin: Look for [voice-call] prefix
  • Call events: openclaw voicecall tail

Useful Debug Commands

# Check webhook endpoint
curl -X POST https://your-url/voice/webhook -H "Content-Type: application/json" -d '{}'

# Verify Tailscale Funnel
tailscale funnel status

# Test TTS locally
openclaw tts "Test message" --provider openai --voice alloy

# Check plugin status
openclaw channels status --deep

# Follow gateway logs
tail -f ~/.openclaw/logs/gateway.log | grep voice-call

Common Issues & Solutions

Issue Symptom Solution
Webhook not reached Call initiated but no events Check public URL, verify firewall
Signature verification fails Events rejected in logs Ensure publicUrl matches provider config
No audio on call Call connects but silent Check TTS provider config/credits
getUserMedia error Browser blocks microphone Use HTTPS, check permissions
Certificate warning Browser shows "Not Secure" Install mkcert CA on device
Call stuck in "ringing" Status never updates Webhook events not reaching gateway

Performance Benchmarks

Target Latencies (P95)

Metric Target Measurement
Call initiation (API) < 2s Time to receive callId
Webhook processing < 200ms Event received → processed
TTS generation < 3s Text → audio ready
STT transcription < 1s Speech end → transcript
WebSocket handshake < 500ms Upgrade request → open
End-to-end (user speech → AI response start) < 5s In conversation mode

Load Testing (Optional)

# Concurrent call simulation
for i in {1..10}; do
  openclaw voicecall call --to "+155500${i}0000" --message "Load test $i" &
done
wait

Monitor:

  • Gateway CPU/memory usage
  • Webhook response times
  • Provider rate limits

Acceptance Criteria

  • All smoke tests pass in CI
  • Manual test cases TC-001 through TC-009 verified
  • HTTPS setup documented and tested with mkcert
  • getUserMedia works on localhost and LAN (with HTTPS)
  • Call latencies within target benchmarks
  • Error messages are actionable
  • Logs provide sufficient debugging info
  • Documentation updated with test findings

References