GPT-6 Luna

OpenAI

GPT-6 Luna stays close to its records. When a customer gets an order detail wrong, it corrects them, and it doesn't guess when its policy has no answer.

It doesn't always finish the job. It tends to ask one more question instead of filing a return or booking a sales call, and it backs down when a customer claims a looser policy. It scores 87 on quality, its median reply takes 4.8 seconds, and each conversation costs $0.0039.

87 /100
#1 of 15 by quality

Level with the runner-up

  • Tool use #5 83
  • Task completion #7 85
  • Context retention #6 80
  • Grounding #7 87
  • Safety #4 93
  • Hallucinations #1 90

Running it

Median reply

4.8s #12 of 15

+3.6s vs the fastest

p95 reply

13.9s

1 reply in 20 is slower

Cost per conversation

$0.0039 #2 of 15

1.1x the cheapest

Consistency

74%

of repeat runs ended the same way

Output per turn

263

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Safe under prompt attack: it refuses to read out an instruction planted in a product record.
  • Grounded: it won't guess a shipping time when its policy gives none.

Watch-outs

  • Acts before confirming: it starts booking a call for a lead who hasn't asked for one.
  • Ungrounded: it lets a customer's overstated warranty stand instead of citing the published term.

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

15 models · 6 scenarios · 46 tests · 3 repetitions · 2070 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by OpenAI. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.