Private LLM Inference - Long
POST https://api.erb-llm.com/v1/long
Privacy-first long-form inference: 8,192-token output cap on a local open-weight model on dedicated hardware - prompts never reach OpenAI/Anthropic or any third-party cloud. OpenAI-compatible chat completions (messages array in, chat.completion JSON out) for summarization, drafting, extraction, translation, code, multi-paragraph answers. Pay per call in USDC on Base (x402), no API key or account. Shorter work? Use the cheaper quick tier (4,096 tokens).
How to buy
This endpoint takes pay-per-call payments over x402. Call it once without paying to see its live price; any x402 client then pays and retries automatically.
curl -i -X POST "https://api.erb-llm.com/v1/long"
# HTTP/1.1 402 Payment Required: price, asset and payTo are in the responseRecent checks
| Checked (UTC) | Result | Status | Time | Live price |
|---|---|---|---|---|
| 2026-10-11 08:15 | Valid 402 | 402 | 1456 ms | $0.05 |
| 2026-10-11 03:45 | Valid 402 | 402 | 947 ms | $0.05 |
| 2026-10-10 23:30 | Valid 402 | 402 | 907 ms | $0.05 |
| 2026-10-10 18:00 | Valid 402 | 402 | 893 ms | $0.05 |
Receipts
No test purchases yet. Graded test purchases are next on our list; until then every number here comes from unpaid checks.
What a check proves