Are Opus 5.5 and GPT-6 Sol Good at Korean?

Written by

in

I recently started a project which benchmarks AI model’s Korean capabilities, hoping it can help which models to choose when doing Korean test.

And of course when I saw the release of the new models today I wanted to test them!

So how does the results look? Lets get straight to the point.

Bencmark results of AI models

Opus 5.5 is the new leader by a tiny margin beating GPT-6 Astra. And GPT-6 Sol is basicly the same as 5.6 Sol. Lets dive deeper.

LevelOpus 5.5GPT-6 Sol
1181/182180/182
2209/220207/220
3243/298216/298

In level 1 and 2 both are basically equal. Level 3 made the difference. Opus was better at solving harder problems.

The ouput itself was also interesting. Opus outputed it’s thought process and Sol just outputed the answer.

And Opus does burn a lot of tokens😭

LevelGPT-5.6 SolGPT-6 Sol
1177/182180/182
2208/220207/220
3214/298216/298

Compared with 5.6 there is basically no difference.

Token effeciency did improve though. Costed a bit less than 5.6 to run the benchmark.

Final thoughts

Frontier model’s Korean capabilities are improving fast. Might need to find a new benchmark soon!

Project website: https://friendly-blini-fe2c97.netlify.app/

Project repo: https://github.com/ahndevdotcom/ko-bm

Benchmark: https://knlp.snu.ac.kr/research/benchmarks

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *