Private LLM Inference - XL
POST https://api.erb-llm.com/v1/extended
Privacy-first extended inference: the highest output cap (32,768 tokens) on a local open-weight model on dedicated hardware - prompts never reach OpenAI/Anthropic or any third-party cloud. OpenAI-compatible chat completions for long-form writing, reports, full document drafts, large code outputs (messages array in, chat.completion JSON out). Pay per call in USDC on Base (x402), no API key or account. Pick the cheaper quick/long tiers for shorter work.
How to buy
This endpoint takes pay-per-call payments over x402. Call it once without paying to see its live price; any x402 client then pays and retries automatically.
curl -i -X POST "https://api.erb-llm.com/v1/extended"
# HTTP/1.1 402 Payment Required: price, asset and payTo are in the responseRecent checks
| Checked (UTC) | Result | Status | Time | Live price |
|---|---|---|---|---|
| 2026-10-11 11:01 | Valid 402 | 402 | 2540 ms | $0.2 |
| 2026-10-11 06:30 | Valid 402 | 402 | 963 ms | $0.2 |
| 2026-10-11 02:15 | Valid 402 | 402 | 1098 ms | $0.2 |
| 2026-10-10 21:45 | Valid 402 | 402 | 1202 ms | $0.2 |
Receipts
No test purchases yet. Graded test purchases are next on our list; until then every number here comes from unpaid checks.
What a check proves