Are Opus 5.5 and GPT-6 Sol Good at Korean?

I recently started a project which benchmarks AI model’s Korean capabilities, hoping it can help which models to choose when doing Korean test.

And of course when I saw the release of the new models today I wanted to test them!

So how does the results look? Let’s get straight to the point.

Bencmark results of AI models

Opus 5.5 is the new leader by a tiny margin beating GPT-6 Astra. And GPT-6 Sol is basically the same as 5.6 Sol. Lets dive deeper.

LevelOpus 5.5GPT-6 Sol
1181/182180/182
2209/220207/220
3243/298216/298

In level 1 and 2 both are basically equal. Level 3 made the difference. Opus was better at solving harder problems.

The ouput itself was also interesting. Opus outtputed it’s thought process and Sol just outtputed the answer.

And Opus does burn a lot of tokens😭

LevelGPT-5.6 SolGPT-6 Sol
1177/182180/182
2208/220207/220
3214/298216/298

Compared with 5.6 there is basically no difference.

Token efficiency did improve though. Costed a bit less than 5.6 to run the benchmark.

Final thoughts

Frontier model’s Korean capabilities are improving fast. Might need to find a new benchmark soon!

Project website: https://friendly-blini-fe2c97.netlify.app/

Project repo: https://github.com/ahndevdotcom/ko-bm

Benchmark: https://knlp.snu.ac.kr/research/benchmarks

Published
Categorized as AI

Leave a comment

Your email address will not be published. Required fields are marked *