Recent complete observations
- Loading current observations…
CODEX MODEL + EFFORT SELECTOR
Describe the task. Codex Juice recommends Astra, Sol, Terra or Luna and the smallest reasoning effort likely to finish it well.
Back to Codex RadarTASK → MODEL → EFFORT
Three task signals produce a transparent starting point. If verification fails, follow the fallback instead of guessing from Juice values.
These are starting points, not capability rankings. Verification difficulty and quality priority can move a task upward.
| Task | Start with | Effort |
|---|---|---|
| Quick and easy to check | Luna | Low |
| Routine and scoped | Terra | Medium |
| Complex repository work | Sol | High |
| Ambiguous or high-risk | Astra | XHigh |
30-SECOND GUIDE
Start with task shape: use less effort when the path and check are obvious, and more when ambiguity or review risk is high. This guide uses OpenAI's quality-versus-speed direction plus external task results—not Juice values.
TECHNICAL MONITOR
The live six-effort observations below monitor collection health and configuration changes. Juice is not a capability score; size and movement are neutral runtime fingerprints.
unavailable
| effort | current | previous | delta | verification |
|---|---|---|---|---|
| Low | — | — | — | unavailable |
| Medium | — | — | — | unavailable |
| High | — | — | — | unavailable |
| XHigh | — | — | — | unavailable |
| Max | — | — | — | unavailable |
| Ultra | — | — | — | unavailable |
Loading current observations…
Choose an effort to inspect confirmed values across retained sweeps.
A restricted, ephemeral Codex CLI session observes one effort at a time. A new value is published only after a second observation from a fresh session matches it. Full sweeps are invalidated if the model identifier changes mid-run.
Juice is a non-official, self-reported runtime fingerprint—not actual reasoning-token use or a quality benchmark. A higher value, lower value, or change in either direction cannot be interpreted as stronger or weaker coding ability.
OpenAI provides only directional quality-versus-speed guidance for reasoning effort and says not every task benefits equally from more reasoning. Read OpenAI’s reasoning-effort guidance
NixBench currently reports GPT-5.6 Sol across Low, Medium, High, XHigh and Max on a 29-task coding corpus, with five recorded trials per effort. This is directional evidence from a different benchmark, not a Codex Juice capability score. NixBench has published no GPT-6 Astra rows yet (checked 2026-09-06).
| effort | mean passed | 95% interval | seconds/task | trials | timeouts |
|---|---|---|---|---|---|
| Low | 23.6/29 | 22.2–25.0 | 27.4s | 5 | 0 |
| Medium | 24.0/29 | 22.5–25.5 | 38.5s | 5 | 0 |
| High | 23.6/29 | 22.9–24.3 | 47.4s | 5 | 0 |
| XHigh | 24.0/29 | 23.1–24.9 | 60.7s | 5 | 0 |
| Max | 23.2/29 | 20.1–26.3 | 81.4s | 5 | 3 |
Limits: NixBench uses its own harness, corpus and objective hidden shell evaluators; these results are not Codex Juice readings. Max has a public row here. Ultra is a product mode with no comparable public performance row found in this sweep.
Static snapshot · verified 2026-08-03
Open the live NixBench results Read the NixBench method Open the NixBench repository Compare the DeepSWE live board Read OpenAI’s GPT-5.6 context