Pay by the second.
Not by the call.
Feneri bills time-based. You pay for the seconds of media the engine actually processes. Every published rate is at least 4x our measured compute cost. Real-time, on-demand, GPU, and storage are priced separately.
Cost a session in five seconds.
Drag the sliders to size your workload. We blend real-time, on-demand, GPU, and storage into a per-month estimate, the same costCents that bills you.
Unit pricing by use
Mix and match. Most teams blend real-time and on-demand depending on workload. Sovereign deployments lock in pre-paid rates.
Live face + voice on CPU. About 2 fps emotion and gaze, about one audio update per second. For agents, live calls, and games.
Async sitting: DeepFace, wav2vec, and Whisper large-v3 on a T4. Priced above realtime because the GPU path is the expensive one.
Voice-only path for telephony, IVR, customer ops. No vision compute. Same wav2vec pipe as the fused stream.
Call recordings, voicemail, podcast / interview analysis. Vocal affect without the Whisper GPU.
Linguistic sentiment and intent. For chat, email, transcripts that arrive without media.
Whisper and peak surges. Billed by the second after a 60s minimum. 4x our T4 box cost.
Affective state, transcripts, narrative reports. Raw video and audio are always deleted within 60 seconds.
Threshold and transition events to your endpoint. Beyond 1M: $0.20 per 1K. Retries free.
Free within a region. Cross-region only when you opt into multi-region replication.
Plans
Commit annually for lower unit prices and SLAs. Or stay pay-as-you-go and never touch a contract.
Sandbox
- 60 min real-time or 200 min on-demand
- 1 GB storage · 30-day retention
- One region · shared compute pool
- Playground + Studio + MCP
- Community support
Build
- Same unit prices as the rate card
- 3 regions · auto-routing
- WebSocket streaming
- All 6 SDKs + MCP + Studio
- Outbound webhooks · 1M free
- Email support · 24h
Production
- Volume pricing that never drops below 4x cost
- All 14 regions · auto failover
- Pinned Whisper T4 · shorter cold start
- 99.97% uptime SLA
- Dedicated success manager
- SOC 2 & ISO reports
Sovereign
- Bring-your-own-GPU or hosted
- Custom fine-tunes on your data
- HIPAA · PDPL · ISO 42001
- Private VPC or on-prem
- Named technical partner
- 24/7 incident SLA
Compare plans
| Feature | Sandbox | Build | Production | Sovereign |
|---|---|---|---|---|
| Unit pricing | ||||
| Real-time streaming | 60 min free | $0.036 / min | $0.034 / min | Negotiated ≥ 4x |
| On-demand inference | 200 min free | $0.048 / min | $0.048 / min | Negotiated ≥ 4x |
| Audio-only realtime | incl. | $0.018 / min | $0.017 / min | Negotiated ≥ 4x |
| Audio-only on-demand | incl. | $0.012 / min | $0.012 / min | Negotiated ≥ 4x |
| Text-only | incl. | $0.0008 / 1K tok | $0.0008 / 1K tok | Negotiated |
| Dedicated GPU-hour | — | $2.90 | $2.90 | BYO or flat-rate |
| Storage · GB-month | 1 GB incl. | 10 GB · then $0.04 | 100 GB · then $0.04 | Self-hosted or flat |
| Performance | ||||
| Real-time latency | ~15 s cold | ~0.5 s video · ~0.8 s audio | ~0.5 s video · ~0.8 s audio | tuned |
| Whisper GPU | shared T4 | shared T4 | pinned T4 | reserved or BYO |
| Stream cold start | ~15 s | ~15 s | pinned · short | none |
| Surfaces & ops | ||||
| API + SDKs | ✓ | ✓ | ✓ | ✓ |
| MCP server | ✓ | ✓ | ✓ | ✓ |
| Studio (upload + live) | ✓ · 90s | ✓ | ✓ · clinical | ✓ · clinical |
| Outbound webhooks | — | 1M free | 10M free | Unlimited |
| Regions | 1 | 3 | 14 | Any |
| SLA | — | — | 99.97% | 99.99% |
| Compliance | ||||
| SOC 2 / ISO 27001 | — | — | ✓ | ✓ |
| HIPAA + BAA | — | — | ✓ | ✓ |
| KSA PDPL / GDPR residency | ✓ | ✓ | ✓ | in-residence |
| Custom fine-tunes | — | — | — | ✓ |
| Support | ||||
| Channel | Community | Email · 24h | Success manager | Named partner · 24/7 |
Questions.
Why billed by the second, not by the call?
An "inference" is a fuzzy unit. A 30-second clip and an hour-long session would both count as one call. We bill the actual seconds of media processed, so a 30-second clip is about $0.018 of real-time and an hour is $2.16. Honest, predictable, and matched to what we spend on compute.
Why does on-demand cost more than real-time?
Live face and voice run on CPU (~2 fps video, ~1 audio update per second). A sitting also runs Whisper large-v3 on a T4. That GPU minute is the expensive part, so on-demand is $0.048 and realtime is $0.036. Every published rate is at least 4x our measured Modal cost.
What counts toward storage?
Only the anonymized state metadata: numerical signals, transcripts, and narrative reports. Raw video and audio are always deleted within 60 seconds of inference, on every plan. We never bill for storing media.
Can I bring my own GPU capacity?
Yes, on Sovereign plans. We deploy the engine into your VPC (AWS, GCP, Azure, or bare metal) and bill a flat capacity license. Custom fine-tunes use your reserved capacity for training.
Do you offer non-profit or research discounts?
Yes, down to the 4x cost floor, never below it. Email research@feneri.xyz with your project.
Where can I see real-time usage?
Your dashboard shows minute-by-minute spend across every unit (real-time, on-demand, GPU, storage, webhooks), with hard caps per project and threshold alerts you set.