I recently started a project which benchmarks AI model’s Korean capabilities, hoping it can help which models to choose when doing Korean test.
And of course when I saw the release of the new models today I wanted to test them!
So how does the results look? Lets get straight to the point.

Opus 5.5 is the new leader by a tiny margin beating GPT-6 Astra. And GPT-6 Sol is basicly the same as 5.6 Sol. Lets dive deeper.
| Level | Opus 5.5 | GPT-6 Sol |
| 1 | 181/182 | 180/182 |
| 2 | 209/220 | 207/220 |
| 3 | 243/298 | 216/298 |
In level 1 and 2 both are basically equal. Level 3 made the difference. Opus was better at solving harder problems.
The ouput itself was also interesting. Opus outputed it’s thought process and Sol just outputed the answer.
And Opus does burn a lot of tokens😭
| Level | GPT-5.6 Sol | GPT-6 Sol |
| 1 | 177/182 | 180/182 |
| 2 | 208/220 | 207/220 |
| 3 | 214/298 | 216/298 |
Compared with 5.6 there is basically no difference.
Token effeciency did improve though. Costed a bit less than 5.6 to run the benchmark.
Final thoughts
Frontier model’s Korean capabilities are improving fast. Might need to find a new benchmark soon!
Project website: https://friendly-blini-fe2c97.netlify.app/
Project repo: https://github.com/ahndevdotcom/ko-bm
Benchmark: https://knlp.snu.ac.kr/research/benchmarks