Make modelslike it's nothing

The custom model factory for cheaper, better, faster AI.

npm install belvedir --save

Collect Your Raw Data

Install the package and Belvedir records every call your agent makes, including tool use and the final reply. You keep all of it.

Search⌘K

Conversations

Every session your agents handled, reconstructed from their traces. Click a row to read the full exchange.
ConversationSpansModelTokensStarted
Refund request for order #4821312claude-sonnet-43,204/61212s ago
session sess_9f2a41c0
user
Hi, I ordered a jacket two weeks ago (order #48213) and it arrived with a broken zipper. Can I get a refund?
▸ tool call · lookup_order
{ "order_id": "48213" }
▸ tool result · lookup_order
{ "delivered": "2026-09-05", "region": "EU", "total": 189.00, "return_window_days": 30 }
assistant

Sorry about the zipper. Your order was delivered on September 5 to an EU address, so it is inside the 30-day return window. I can issue a full refund of €189.00 to your original payment method, or send a replacement at no charge. Which would you prefer?

Package tracking shows no movement9claude-sonnet-42,118/38841s ago
Cancel subscription before renewal7gpt-4.11,876/3011m ago

Clean and Build Datasets

Traces are grouped by task and stripped of noise. You can read through the result before you train on it.

Search⌘K

Task Groups

Recurring task types, clustered from your traffic. A task group with enough clean examples can become a training environment.
Search task groupsAll API keysLast 30 daysCleaning pass 2h ago · 2 merges, 1 rename
Refund eligibilityWhether a specific order qualifies for a refund, and under which policy clause.
1,840 clean examples · ready71% success48 tasks
Order status lookupWhere an order is, when it will arrive, and why tracking has not moved.
1,412 clean examples · ready94% success41 tasks
Policy citationQuoting the exact return, warranty, or shipping policy that applies.
960 clean examples · ready62% success33 tasks
Escalation routingDeciding when a case needs a human and which team should get it.
412 clean examples88% success29 tasks
Invoice parsingPulling totals, dates, and line items out of uploaded invoices.
388 clean examples79% success24 tasks
Address changesUpdating a shipping address on an order that has not left the warehouse.
301 clean examples91% success22 tasks
Tone rewriteSoftening or formalizing a drafted reply before it is sent.
204 clean examples96% success19 tasks

Train on What Matters

Fine-tune an open model on that dataset. Each run is saved with its config and data so you can come back to it later.

Search⌘K

Training

support-triage-ft.4

completedsftautofrom Qwen/Qwen3.8-27B·9/7/2026·produced support-triage-ft.4

Examples

1,840

6,212 turns

Eval loss

1.92 → 0.61

base → trained

Train loss

0.58

Duration

42 min

$3.120 infra

Training loss

Training loss per step

2.460
1.503
0.547
training starttraining end

Dataset

1,840 examples · 6,212 turns · 1,656 train / 184 eval · refund-eligibility, policy-citation

Data through 9/6/2026

Configuration

Epochs

3

Split

1656 train / 184 eval

Training tokens

4,812,930

backend modal · lr 0.00002 · lora r 32 · seq 4096 · exit 0

Prove It Got Better

Each model is scored on your own evals next to the base model and a frontier model, so you can see what changed before anyone uses it.

Search⌘K

Back to evals

support-bench

Run
Tasks from Refund eligibility · cloud sandbox

Score History

support-triageorder-statusgpt-4.1referenceQwen/Qwen3.8-27Breference

Runs

support-triage-ft.40.6122d ago
order-status-ft.20.5189d ago
support-triage-ft.40.59011d ago
gpt-4.10.55211d ago
support-triage-ft.3–3w ago

Serve Your Own Model

Point your existing calls at the new model, or download the weights and run them wherever you like.

Search⌘K

Models

support-triage-ft.4

Fine-tuned from Qwen/Qwen3.8-27B

ServingQwen/Qwen3.8-27B · 1,840 examplesTrained 9/7/20260.590
PlaygroundRun benchmarkDeploy

Serving on Belvedir cloud inference. Point any OpenAI-compatible client at the endpoint with a project key.

curl https://platform.belvedir.ai/api/v1/route/chat/completions \
  -H "Authorization: Bearer $BELVEDIR_API_KEY" \
  -d '{
    "model": "support-triage-ft.4",
    "messages": [{ "role": "user", "content": "..." }]
  }'

Benchmarks

benchmark_score 0.590eval_loss 0.61base_eval_loss 1.92
0.612support-bench · refund-eligibility2d ago
0.584support-bench · policy-citation2d ago
0.574support-bench · full11d ago

Versions · 2

support-triage-ft.3
Retired8/17/20260.531
support-triage-ft.2
Retired7/29/20260.488

Frequently asked questions

Into your own account. Traces, data, and trained models stay isolated to your workspace and never train anyone else's models.

It turns your production traces into training data, trains on them, and scores each run against your own benchmarks.

No. Training and inference run on our infrastructure. You can export the weights and self-host at any time.

Others start from a dataset you have already built. Belvedir starts from your production traffic and keeps training as new traffic comes in.

Install the package and sign in to create a workspace. Larger teams can get in touch for onboarding.

Cheaper, better, faster AIis one command away

npm install belvedir --save