Voice & Multimodal
Vapi
Developer platform for voice AI agents over phone, web, and MCP
Category
Platforms that let agents perceive, generate, or interact through voice, audio, images, video, and other modalities.
Editorial scope version 1.0 · effective 3 September 2026
Category listing
Definition
This category covers agent-first interaction layers where real-time speech or another non-text modality is central to the agent’s operation or user experience.
Includes real-time voice and other non-text agent interaction layers that manage modalities, transports, turn-taking, interruptions, state, media events, or handoffs.
Scope
Selection guide
Category evidence
The GPT-Realtime-2 documentation describes speech-to-speech agent workflows with text, image, and audio inputs, audio output, and tool use.
accessed 2026-09-03Comparable facts
Unknown values are shown explicitly rather than inferred from marketing copy. Open a profile to inspect its claim-level sources.
| Tool | Classification | Pricing | Interfaces | Deployment | Verification |
|---|---|---|---|---|---|
| Vapi | Agent-native | Paid | REST API, web dashboard, Web SDK, TypeScript server SDK, Python server SDK, command-line interface, MCP server, webhooks | hosted | documentation reviewed |