Animated WaffleDocsLog in

Public beta

Embed an Agent in your application

Connect a published Agent to a browser without exposing realtime transport, provider configuration, or permanent API keys. This guide assumes nothing beyond a published Agent.

Two packages

Choose your integration

Pick the smallest package that matches who owns rendering.

@animated-waffle/client

Framework-independent lifecycle, microphone input, text input, state, and transcripts. Your app owns the UI.

@animated-waffle/react

The client API plus React state, Agent audio, and the published Avatar renderer.

Before you code

Three things you need

  1. A published Agent

    Build and publish it in the Dashboard first — the Agents guide covers every step. Session tokens are only minted for published Agents.

  2. A workspace API key

    An Owner or Admin creates it under Integrations in the Dashboard sidebar (direct link: /operator/keys): enter a Label, select Create key, and copy the awp_… value immediately — it is shown exactly once. Full walkthrough in the API reference.

  3. The Agent’s id

    Your backend calls GET /v1/agents with the key and selects your Agent by its slug; the response’s id is the stable UUID used everywhere below.

Public beta

Install from npm

Terminal
npm install @animated-waffle/client

# React applications
npm install @animated-waffle/react

Both packages are published on npm. @animated-waffle/react already includes the client — install it alone for React applications.

The default managed-session identity and sessionContext APIs were introduced in version 0.3.1. Use @latest for the current registry release.

Trust boundary

Mint session tokens on your backend

Permanent workspace keys belong on your backend. The browser asks your backend for one short-lived Agent token per connection; the SDK does the rest.

server.ts · your token endpoint
// Your backend. The awp_ key comes from Dashboard → Integrations
// and never leaves the server.
app.post('/api/animated-waffle/session-token', async (req, res) => {
  const user = await authenticateYourUser(req)
  if (!user) return res.status(401).end()

  const response = await fetch(
    `https://animated-waffle.narya.ai/v1/agents/${AGENT_ID}/session-tokens`,
    {
      method: 'POST',
      headers: {
        Authorization: `Bearer ${process.env.NARYA_API_KEY}`,
        'Content-Type': 'application/json',
      },
      // end_user_id: required when the Agent has memory or Calendar
      // enabled; must be a stable UUID for this user in your system.
      body: JSON.stringify({ end_user_id: user.stableUuid }),
    },
  )
  if (!response.ok) return res.status(502).json(await response.json())

  const { token } = await response.json()
  res.json({ token })
})
What Narya returnsHTTP 201 with { token, expires_at, agent_id, capabilities }. The token is valid for 15 minutes and starts exactly one Agent conversation.
Authorize your own user firstBefore minting, authenticate the user, verify they may talk to this Agent, and apply your own rate, duration, concurrency, and spend limits.
end_user_idOptional unless the Agent has conversation memory, documents, or Calendar enabled — then it must be a stable UUID for this user from your identity system. Send {} otherwise.

Framework independent

Connect with JavaScript

Use the client package when your application already owns rendering and controls.

client.ts
import { WaffleClient } from '@animated-waffle/client'

const agentId = '22222222-2222-4222-8222-222222222222'
const client = new WaffleClient({
  agentId,
  sessionToken: async () => {
    const response = await fetch('/api/animated-waffle/session-token', {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({ agentId }),
    })
    if (!response.ok) throw new Error('Could not start a session')
    return (await response.json()).token
  },
})

client.on('transcript', ({ id, role, text, final }) => {
  if (final) console.log(id, role, text)
})
client.on('declined', ({ reason }) => {
  settleCurrentTurn({ spoken: false, reason })
})
client.on('connectionState', (state) => console.log('connection:', state))

await client.connect({ inputMode: 'push-to-talk' })
Option / memberWhat it does
agentIdThe published Agent’s UUID from GET /v1/agents. Must be a UUID, not a slug.
sessionTokenA function (sync or async) returning a fresh token from your backend. Called on each connect. Alternatively pass startSession to start sessions through your own authenticated backend; the two options are mutually exclusive.
sessionContextDefault sessionToken starter only: optional { userName?, timeZone? }. Nonblank values are sent exactly as provided; blank values are omitted. Mutually exclusive with startSession, so token-path consumers do not recreate the raw managed-session request.
apiOriginOptional. Point at https://animated-waffle-dev.narya.ai for the development environment; defaults to production.
inputAudioEnabledOptional. Set false for text-triggered, output-only applications; the SDK never requests the microphone and PTT controls fail closed with input_audio_disabled. Defaults to true and stays off until the Agent is ready.
audioPlaybackOptional, default true: the client plays the Agent’s voice through its own audio element. Set false to play the remoteAudioTrack yourself — an app routing the voice through its own element or Web Audio graph must turn this off, or the voice is heard twice.
connect({ inputMode })‘push-to-talk’ (default) or ‘automatic’ for an open microphone with voice detection. Resolves with the session.
sendText(text, { audioResponse })Sends a typed message into the conversation. audioResponse: false makes only that turn text-only.
mutedRead/write. Silences the Agent’s voice while the session, transcript, and Avatar performance carry on. No effect when audioPlayback is false.
startPushToTalk() / commitPushToTalk() / cancelPushToTalk()Hold-to-talk controls. Throw unless connected with inputMode ‘push-to-talk’.
on(‘connectionState’ | ‘session’ | ‘transcriptProgress’ | ‘transcript’ | ‘declined’ | ‘agentSpeaking’ | ‘remoteAudioTrack’ | ‘error’, fn)Event subscriptions; each on() call returns its unsubscribe function. declined carries { reason } for deliberate silence. transcriptProgress is ephemeral ASR/playout caption state keyed by runtime segmentId. transcript events are completed persisted rows with the same required id as history and currently final=true.
disconnect()Ends the session and releases the microphone.

The managed SDK stays provider-neutral: 0.3 removes provider-specific overrides such as openaiModel; use the separate OpenAI Realtime page when that provider-specific path is required.

Media-vendor credentials, room credentials, and Avatar asset URLs stay inside the SDK — your code never handles them.

React

Connect and render with useWaffle

Use the React package when Aniwaffle should own Agent audio and Avatar playback.

Agent.tsx
import { AvatarStage, useWaffle } from '@animated-waffle/react'

export function Agent() {
  const agentId = '22222222-2222-4222-8222-222222222222'
  const waffle = useWaffle({
    agentId,
    sessionToken: () => fetchSessionToken(agentId),
  })

  return (
    <>
      {/* Session churn rebinds signals without removing the visible stage */}
      <AvatarStage {...waffle.avatar} style={{ width: 480, height: 480 }} />
      {waffle.avatarLoad.status === 'downloading' && (
        <progress value={waffle.avatarLoad.progress ?? undefined} />
      )}
      <button
        disabled={waffle.connectionState !== 'disconnected'}
        onClick={() => void waffle.connect()}
      >
        Start conversation
      </button>
      <button onClick={() => void waffle.disconnect()}>End</button>
    </>
  )
}

The hook returns the app-facing client state — connectionState, session, transcriptProgress, transcript, decline (the latest deliberate silence), agentSpeaking, muted/setMuted, performance (the cue the Avatar is performing right now), error, connect, disconnect, sendText, and the push-to-talk controls — plus an opaque avatar binding. Spread it onto a sized AvatarStage and the published character renders itself without exposing the remote audio track. The returned action functions keep stable identity across renders. avatarLoad reports first-load progress (downloading → verifying → parsing → ready with byte counts); repeat visits come from the browser cache after hash verification, so loading is near-instant.

Disconnect and reconnect only rebind session signals. The visible renderer is destroyed when AvatarStage unmounts.

When it fails

Errors you will actually see

FailureCause and fix
Your token endpoint gets 409 “Publish the agent before requesting a session token.”The Agent is still a draft. Open the Agent editor and select Publish in the Dashboard.
Your token endpoint gets 400 memory_identity_required, documents_identity_required, or calendar_identity_requiredThe Agent has conversation memory, documents, or Calendar enabled but the request had no end_user_id. Send a stable UUID for the user.
Your token endpoint gets 401 unauthorizedThe awp_ key is missing, wrong, or revoked — check the Integrations page and your secret store.
The SDK throws ‘agentId must be a UUID.’A slug was passed as agentId. Resolve the slug to the id through GET /v1/agents on your backend.
connect() rejects with an expired-token errorTokens live 15 minutes and the sessionToken callback returned a cached one. Mint a fresh token on every call.
Push-to-talk methods throwThe client was connected with inputMode ‘automatic’. Reconnect with ‘push-to-talk’ or rely on voice detection.

Before launch

Production checklist

  • Publish and test the Agent in Live conversation before wiring the SDK.
  • Keep the awp_ workspace key only on your backend.
  • Bind end_user_id to a stable user in your own product.
  • Apply authorization, rate, concurrency, duration, and spend limits before minting tokens.
  • Handle disconnected and error states in the application UI.
  • Size AvatarStage and show avatarLoad progress on first load.

SDK migration

Upgrade from SDK 0.3 to 0.4

Version 0.4.0 is a breaking release of the client, server, and React packages. Character Card V3 is the complete character configuration. Replace these SDK fields before upgrading:

Previous fieldReplacement
Agent displayNamecharacterCard.data.name
Agent management descriptioncharacterCard.data.creator_notes
Agent personacharacterCard.data.description
Agent instructionscharacterCard.data.system_prompt
Session overrides.persona / overrides.systemPromptoverrides.characterCard with the corresponding description / system prompt
Authoring updatedPersona / updatedPromptupdatedCharacterCard
Authoring base revision persona / systemPromptcharacterCard
Authoring promptHashRemoved; candidate operations continue to use candidateHash

Card writes replace the complete card. Preserve its other fields when editing one part of an existing Agent:

ts
const agent = await waffle.getAgent(agentId)
await waffle.updateAgent(agentId, {
  characterCard: {
    ...agent.characterCard,
    data: { ...agent.characterCard.data, system_prompt: 'Keep replies concise.' },
  },
})

Push-to-talk controls return promises so callers can await queued controls. Update custom WaffleSessionClient and UseWaffleResult implementations to match these signatures. WaffleSessionClient also exposes microphoneTrack, setInputMode, setMicrophoneEnabled, and interrupt. Existing calls that ignore the built-in client's return value still work; control failures continue to emit typed error events. Application sessions require an API and worker deployment advertising protocol 2 before the new consumer is enabled.

Backend extensions

Application-owned tools and conversation state

Client and server SDK 0.4.0 support backend-owned application sessions. The backend supplies instructions, history, and JSON Schema tools; it joins a separate data-only connection to handle input routing and tool execution. Permanent workspace keys and the control grant stay on the backend. Browser session tokens and Viewer Companion credentials cannot start this session type. The application owns the tool list: published memory, calendar, document, web-search, and stay-quiet capabilities are disabled for these sessions.

ts
import { WaffleServer } from '@animated-waffle/server'
import { WaffleApplicationSession } from '@animated-waffle/server/realtime'

const api = new WaffleServer({ apiKey: process.env.ANIMATED_WAFFLE_API_KEY! })
const started = await api.startApplicationSession({
  agentId,
  chatSessionId: crypto.randomUUID(),
  requestId: crypto.randomUUID(),
  application: {
    instructions: 'Help the user search their own records.',
    history: [],
    tools: [{
      name: 'search', description: 'Search records.',
      parameters: { type: 'object', properties: {}, additionalProperties: false },
      cancellable: true, onDuplicate: 'reject',
    }],
  },
})
const control = new WaffleApplicationSession(started.control, {
  async handleRequest(request, { signal, update }) {
    switch (request.type) {
      case 'input': return { action: 'reply', text: request.text }
      case 'tool':
        await update({ status: 'Searching' })
        return searchAuthorizedRecords({ signal })
      case 'transcript': await saveTranscript(request); return null
      case 'prepare_reply': return { instructions: await buildCurrentInstructions(request.id) }
    }
  },
  onError: reportError,
})
await control.connect()
// Keep control alive with the application's session owner.
return Response.json(started.session)

Install the optional Node transport with npm install @animated-waffle/server @livekit/rtc-node@0.13.33. Importing the server SDK's main entry does not load the native transport. The backend must authenticate its caller and authorize all tools before creating a session. The control token expires after five minutes for joining, is room-scoped, and cannot publish or subscribe to audio/video. It can send data to its worker. Only started.session is returned to the browser:

ts
const client = new WaffleClient({
  application: true,
  startSession: () => fetch('/api/voice/session', { method: 'POST' }).then(r => r.json()),
})
await client.connect()

The backend callback receives:

  • input: finalized text, id, and modality (audio or text). Return { action: 'handled' } or { action: 'reply', text, instructions? } before the foreground model responds.
  • tool: a declared name, arguments, and tool-call id. Validate arguments and authorize the operation. Await update(value) to report JSON progress. Its first update makes the tool non-blocking; subsequent updates and the final result use LiveKit's idle-aware delivery. Progress can trigger conversation.
  • transcript: durable message id, role, and text. Persist idempotently and acknowledge. This direct backend delivery does not rely on the browser.
  • prepare_reply: a spoken update id. Return { instructions } after preparing current context and any visual focus needed before speech.

control.updateContext({ instructions, messages }) replaces instructions and appends any supplied messages without requesting speech. Await each update to preserve application ordering. control.requestReply(id) explicitly queues a spoken update, asks the backend to prepare it when idle, and returns { id, interrupted, completion: 'worker_playout' }. This means the worker's playout completed or was interrupted; it is not a client speaker acknowledgement. Reusing an ID in the same connection does not speak twice. New user input invalidates queued updates. The application owns replay across connections.

Tools default to cancellable: false, onDuplicate: 'allow', and runReply: true. Native duplicate policy (allow, reject, replace, confirm) compares concurrent tool names, not business arguments. Use runReply: false for dispatch tools whose eventual spoken result arrives via requestReply; business idempotency still belongs to the application. control.cancelTool(toolCallId) requests explicit cancellation. After its first progress update, an async tool survives speech interruption. Tools without progress retain normal blocking-tool interruption behavior. Honor the callback signal for connection-bound work; a durable business task needs its own owner and explicit business cancellation, independent of audio and session lifetime.

Browser controls remain client.setInputMode('automatic' | 'push-to-talk'), client.setMicrophoneEnabled(enabled), client.interrupt(), and PTT controls. client.microphoneTrack supports microphone meters. PTT promises settle after the existing control queue; failures also emit error.

The start request has a 256 KiB body limit, with at most 200 history messages and 64 tools. Application text streams allow 256,000 characters per envelope and 128 pending calls. Requests need a result or first tool progress within 60 seconds; acknowledged async tools then share the session lifetime. Responses use type (progress, result, error) and the same UUID requestId in the body and stream attributes. Values must be JSON-serializable; undefined results become null. Late responses are ignored after completion or cancellation. Backend commands use application.command / application.command_result; worker callbacks use application.request / application.result and application.cancel. Only the bound backend identity can send commands/results; only the current Agent can invoke callbacks. The capability handshake requires animated-waffle.application.version=2. Deploy the API and worker before connecting an application client; older workers fail closed.

The runnable application-session example checks slow tools, speech interruption, cancellation, context-only updates, reply completion, and connection cleanup without depending on Canvas.