YGT Labs AIGPT‑5.6 DECISION ATLASTR edition

YGT LABS AI / RESEARCH EDITION · JULY 10, 2026

Three models.
One clear
route.

Choose GPT‑5.6 Sol, Terra or Luna by the real bottleneck: capability, daily delivery or high-volume speed. This English edition keeps the benchmark evidence, local measurements and operating limits readable without hiding the trade-offs.

40
official benchmark rows
5
independent sources
34
isolated local runs
17
Standard / Fast pairs

MODEL ROLES

Capability, balance
and throughput.

These are decision roles, not claims that one model is universally best. The intelligence index is a max-reference comparative index; local speed comes from a single fixed-task run per cell.

Capability ceiling

Sol

58.9intelligence

For architecture, difficult debugging, security-sensitive reasoning and high-cost mistakes.

Fast effective speed
51.8 tok/s
API output list price
$30 / 1M
Highest local mode
Ultra
Daily driver

Terra

55intelligence

For daily coding, research and controlled agent delivery with a strong price-to-capability balance.

Fast effective speed
60.7 tok/s
API output list price
$15 / 1M
Highest local mode
Ultra
Speed and scale

Luna

51.2intelligence

For summaries, classification, first drafts and high-volume helper tasks after quality has been measured.

Fast effective speed
68.4 tok/s
API output list price
$6 / 1M
Highest local mode
Max

REASONING LEVELS

More thinking changes
the resource budget.

Effort can improve how much work a model invests in a task. It is not a universal IQ dial, and published effort curves do not exist for every downstream capability.

  1. 01

    Low

    Short planning budget for quick, bounded work.

    Sol · Terra · Luna
  2. 02

    Medium

    Balanced daily problem-solving budget.

    Sol · Terra · Luna
  3. 04

    XHigh

    Expanded intermediate reasoning for complex work.

    Sol · Terra · Luna
  4. 05

    Max

    Highest standard single-agent budget for difficult problems.

    Sol · Terra · Luna
  5. 06

    Ultra

    An orchestration mode, not a directly comparable single-agent effort level.

    Sol · Terra
YGT Labs AI

YGT LABS AI / APPLIED AI

Do not just select
a model. Build a system.

YGT Labs AI joins product engineering, AI, IoT and integrated automation. The atlas is the model-routing layer: it makes the starting capability, effort and cost boundary explicit before a workflow reaches production.

01

Product engineering

Architecture options, repository discovery, implementation plans and test-first changes. Use Sol when a wrong answer is expensive; start daily delivery with Terra.

02

Agent operations

Triage support, classify work, prepare research briefs and make human approval points explicit. A model is one part of the system, not the process itself.

03

Content, SEO and GEO

Create source-backed expert content, structured data and multilingual quality checks. The goal is useful, reviewable answers—not automated content volume.

04

Research and data quality

Turn scattered inputs into evidence matrices, decision notes and visible uncertainty. Use a stronger route for conflicts, exceptions and final review.

05

Vertical SaaS and automation

Connect assistants to portals, CRM, IoT or operations with business rules, data ownership and measurable outcomes intact.

EVIDENCE, NOT SLOGANS

Every claim carries
its own boundary.

Official benchmark rows, independent signals and local telemetry answer different questions. We keep their source class and limits separate.

Professional

Agents’ Last Exam

Sol
52.7
Terra
50.4
Luna
50.3

Long-horizon professional tasks.

Coding

AA Coding Agent Index

Sol
80.0
Terra
77.4
Luna
74.6

Composite coding-agent index.

Science

GeneBench Pro

Sol
28.7
Terra
23.3
Luna
10.8

Long scientific workflows.

Cybersecurity

SEC-Bench Pro

Sol
71.2 · Ultra 74.3
Terra
57.7
Luna
48.9

Security-focused agent tasks.

Long context

MRCR 512K–1M

Sol
73.8
Terra
72.5
Luna
41.3

Retrieval across 512K–1M context.

Abstract

ARC-AGI-3

Sol
7.78
Terra
0.80
Luna
0.18

Interactive abstract reasoning.

Local test method

  1. Every cell used a separate ephemeral Codex process.
  2. The task was a fixed weighted-interval-scheduling problem.
  3. Standard and Fast ran across every supported effort.
  4. Wall time, token classes, cache and API-equivalent cost were retained.

What this does not prove

Each local cell has n=1. Queueing, cache and sampling introduce noise. Effective token/s is not decoder throughput, and one task is not a general-intelligence test.

Read the methodology report ↗ Open JSON ↗

PRIMARY SOURCES

OFFICIALOpenAI GPT‑5.6 ↗OFFICIALOpenAI API pricing ↗INDEPENDENTDeepSWE ↗INDEPENDENTARC Prize ↗INDEPENDENTArtificial Analysis ↗