A practical result for Korean AI corrections

Which AI should you use to check your Korean?

A test of AI models on 1,418 sentences written by Korean learners: API, Assistants or Local models.

Practical Results Choose based on what matters most to you.
  1. Free AI Assistant Option

    DeepSeek Free, web and app

    DeepSeek v4.1 flash placed second in the benchmark and is likely the model powering DeepSeek chat.

  2. Best paid option

    Grok 4.6 Paid plan, xAI

    Grok 4.6 clearly outperforms every other option - it is paid only however.

  3. Best offline option

    Muse Glimmer Runs offline on your own hardware

    Meta recently released Muse Glimmer which can comfortably compete with the average paid option. At Q4,K,M it sits at 19GB in size.

Version
2026-09-12
Tested
KoLLA v2 Dataset
Code
GitHub Repository
Updated
September 12, 2026

Overview

The best choice depends on your goal. This benchmark can guide you from privacy and independence with a self hosted Muse Glimmer, to quality-price-complexity tradeoff with the cheap Gemini 3.5 Flash Lite.

Strongest result in this test

Grok 4.6

66.9% F0.5

It performed the best on the whole Dataset, balancing the share of mistakes it caught with the share of useful changes it made. Overall it cost around $8.85 to run for 1000 Sentences.

Cheapest hosted option

Range of Cost to run

$0.06 - $18.34 for one full run

The cost of running the benchmark once on a model via the API has a huge range. Google manages to take the bottom and the lead with its very cheap Gemini 3.5 Flash Lite and its very expensive Gemini 3.1 Pro. Cost can scale double with performance: Tokens get more expensive, but the model also produces more reasoning tokens.

The Price-Quality Frontier

7 hosted options

Gemini 3.5 Flash Lite, MiMo V2.5, GLM 5.3 Flash, GPT-5.6 Terra, Claude Sonnet 5, DeepSeek V4.1 Flash, Grok 4.6

Each one gives you more measured quality than every cheaper option tested.

Price vs. correction quality

Each dot is a hosted model. Moving right costs more; moving up means more reliable corrections. Offline models have no API bill and so no place on this axis; the best of them is drawn as a line across the chart.

Correction quality against estimated cost A comparison of correction quality and estimated cost per one thousand sentences. The highlighted line connects the best tested quality and price trade-offs. A dashed horizontal line marks the best model you can run on your own computer, which has no price and so no place on a cost axis. Free on your own computer: Muse Glimmer Gemini 3.5 Flash Lite Claude Sonnet 5 GPT-5.6 Terra DeepSeek V4.1 Flash Grok 4.6 MiMo V2.5 GLM 5.3 Flash Overall quality F0.5 score Estimated cost for 1,000 sentences

The dashed line is the best local model tested. 6 hosted models you pay for score below it.

The highlighted line connects models that offer more quality than every cheaper model tested. It is a guide to trade-offs, not a universal ranking.

See every model

Compare every completed and unfinished test. Select a column heading to sort the results.

Hosted services

Showing 17 runs

Service
Grok 4.6 xAI OpenRouter Default settings 66.9% F0.5 65.2% 74.5% 52.8% 748 sentences $8.85 1,417 1 failed 1,423 2,352,541 tokens total
DeepSeek V4.1 Flash DeepSeek OpenRouter Default settings 64.4% F0.5 62.7% 72.2% 51.6% 728 sentences $2.82 1,412 6 failed 2,028 2,948,491 tokens total
Kimi K3 Moonshot AI OpenRouter Default settings 63.8% F0.5 62.1% 71.7% 52.2% 739 sentences $9.24 1,417 1 failed 622 1,095,091 tokens total
Claude Sonnet 5 Anthropic OpenRouter Default settings 63% F0.5 61.1% 72.5% 49.4% 701 sentences $1.11 1,418 0 failed 99 224,073 tokens total
Claude Opus 5 Anthropic OpenRouter Default settings 60.7% F0.5 58.1% 74.5% 43.1% 611 sentences $5.74 1,418 0 failed 218 392,003 tokens total
Qwen3.8 Max 0902 Alibaba Cloud OpenRouter Default settings 60.4% F0.5 58.2% 71.1% 48% 668 sentences $11.73 1,391 27 failed 985 1,478,068 tokens total
GPT-5.6 Terra OpenAI OpenRouter Default settings 60.3% F0.5 58% 71.9% 44.9% 637 sentences $1.07 1,418 0 failed 50 123,206 tokens total
GLM 5.3 Flash Z.ai OpenRouter Default settings 60% F0.5 58.2% 68.6% 46% 649 sentences $0.86 1,412 6 failed 995 1,468,601 tokens total
Gemini 3.1 Pro Preview Google OpenRouter Default settings 59.5% F0.5 57.2% 70.8% 47.7% 676 sentences $12.94 1,417 1 failed 1,070 1,553,186 tokens total
GLM 5.3 Z.ai OpenRouter Default settings 59.4% F0.5 57.2% 70.5% 43.4% 613 sentences $6.11 1,412 6 failed 1,289 1,883,143 tokens total
Gemini 3.7 Flash Google OpenRouter Default settings 58.8% F0.5 56.1% 73% 41% 582 sentences $2.16 1,418 0 failed 570 844,638 tokens total
MiMo V2.5 Xiaomi OpenRouter Default settings 58.1% F0.5 56.7% 64.6% 45.2% 640 sentences $0.11 1,417 1 failed 297 477,529 tokens total
Gemini 3.5 Flash Lite Google OpenRouter Default settings 57.6% F0.5 55.3% 68.8% 41% 582 sentences $0.04 1,418 0 failed 14 55,906 tokens total
MiniMax M3 MiniMax OpenRouter Default settings 56.1% F0.5 54.6% 63.5% 42.3% 596 sentences $0.75 1,408 10 failed 523 1,008,569 tokens total
GPT-5.6 Sol OpenAI OpenRouter Default settings 56.1% F0.5 53.5% 69.9% 40.8% 579 sentences $1.09 1,418 0 failed 101 196,316 tokens total
Qwen3.8 Flash Alibaba Cloud OpenRouter Default settings 56% F0.5 54.3% 64% 44.8% 621 sentences $0.97 1,386 32 failed 2,092 2,990,041 tokens total
GPT-5.6 Luna OpenAI OpenRouter Default settings 54.5% F0.5 51.8% 69.2% 37.9% 538 sentences $0.16 1,418 0 failed 130 237,568 tokens total

Run on your own computer

Showing 18 runs

These have no hosted-service bill, so the price column is replaced by the weights that were actually loaded. A smaller file is quicker and fits more machines; it is also a compressed version of the model, which costs accuracy.

Muse Glimmer Meta BF16 30B · 59.6 GB 58.5% F0.5 56.6% 67.6% 43.6% 615 sentences 1,411 7 failed 674 1,025,533 tokens total
Muse Glimmer Meta Q4_K_M 28B · 18.2 GB 57.9% F0.5 55.9% 67.4% 41.7% 591 sentences 1,418 0 failed 693 1,057,607 tokens total
Mi:dm 2.0 Base Instruct KT BF16 11.5B · 23.1 GB 56.4% F0.5 57.8% 51.5% 44.2% 627 sentences 1,418 0 failed 12 766,117 tokens total
Kanana 2 30B A3B Instruct Kakao BF16 30B-A3B · 61.3 GB 53.6% F0.5 52.2% 60.1% 40.1% 568 sentences 1,418 0 failed 12 67,148 tokens total
Gemma 4 26B A4B Google Q4_K_M 26B-A4B · 18.0 GB 52.9% F0.5 50.4% 66.1% 37.4% 527 sentences 1,410 8 failed 1,011 1,482,853 tokens total
Gemma 4 26B A4B Google BF16 26B-A4B · 51.6 GB 52% F0.5 49.4% 65.8% 32.2% 457 sentences 1,418 0 failed 15 84,527 tokens total
Gemma 4 12B Google BF16 12B · 23.9 GB 49.1% F0.5 46.4% 64.1% 26.1% 370 sentences 1,418 0 failed 16 85,725 tokens total
Command A+ 05-2026 Cohere FP8 219B · 225.0 GB 46.5% F0.5 44.4% 57.3% 30.7% 435 sentences 1,415 3 failed 359 710,701 tokens total
Qwen3.6 35B A3B Alibaba Cloud FP8 35B-A3B · 37.5 GB 45.4% F0.5 43.4% 55.7% 31.2% 440 sentences 1,412 6 failed 1,682 2,430,795 tokens total
Nemotron 3.5 Lightning 30B A3B NVIDIA NVFP4 30B-A3B · 21.6 GB 43% F0.5 41.8% 48.7% 30.4% 428 sentences 1,410 8 failed 1,676 2,421,700 tokens total
Gemma 4 E4B Google Q4_K_M 7.5B · 6.3 GB 40.7% F0.5 38.2% 54.8% 16.1% 229 sentences 1,418 0 failed 81 172,332 tokens total
GPT-OSS 20B OpenAI MXFP4 20B · 12.1 GB 39.8% F0.5 37.8% 50.7% 20.5% 291 sentences 1,418 0 failed 92 273,327 tokens total
Kanana 2 30B A3B Thinking Kakao Q4_K_M 30B-A3B · 18.6 GB 36.6% F0.5 34.8% 46.6% 19.5% 276 sentences 1,418 0 failed 924 1,364,980 tokens total
Ministral 3 3B Mistral AI Q8_0 3B · 4.5 GB 28% F0.5 26.3% 37.9% 3.7% 52 sentences 1,418 0 failed 16 65,908 tokens total
Ministral 3 14B Reasoning Mistral AI Q4_K_M 14B · 9.1 GB 26.5% F0.5 25.4% 31.5% 9.7% 138 sentences 1,417 1 failed 15 64,947 tokens total
Nemotron 3 Nano 4B NVIDIA Q8_0 4.0B · 4.2 GB 23.6% F0.5 22.6% 28.4% 8.9% 126 sentences 1,416 2 failed 902 1,336,980 tokens total
DeepSeek R1 0528 Qwen3 8B DeepSeek Q4_K_M 8B · 5.0 GB 23.2% F0.5 21.9% 30.3% 4.3% 60 sentences 1,404 14 failed 382 581,964 tokens total
Kanana 2 3B Instruct Kakao Q6_K 3B · 2.9 GB 12.5% F0.5 12.1% 14.1% 0.8% 12 sentences 1,418 0 failed 13 74,979 tokens total

Unfinished tests stay visible for transparency but are left out of the recommendations. Models for your own computer were run through LM Studio or on rented Vast.ai GPUs; their hardware, electricity, and time were not measured.

Output lengthCompletion tokens / sentence is the average generated text per answered sentence. Models used their normal settings, so some spent much more time thinking about the same one-sentence task. That affects both price and speed.

What models actually do

Two models can have similar scores while giving very different advice. Pick a maker and model to see whether it usually gets the sentence right, changes too much, or misses the mistake.

Change nothing at all: 40.5% fully correct (574 of 1,418 sentences)

Fully correctNo FP, no FN
747 52.7%

No extra change and no missed mistake

Changed too muchFP only
310 21.9%

Added a change that was not needed

Missed a mistakeFN only
73 5.1%

Left a needed correction undone

Mixed resultFP + FN
287 20.2%

Changed too much and still missed something

Could not answerFailed
1 0.1%

No answer was returned after retries

How the test works

The details for readers who want to check the work: the sentences used, the question each model received, how answers were judged, and what this test cannot tell us.

The sentence corpus

I used KoLLA_multi-refs.m2 from the KoLLA v2 dataset: 100 learner essays, split into 1,418 sentences. They contain 2,828 human corrections and 3,649 individual edits. All but eight sentences have two human references.

Fingerprint
f08943bedab1b149
MD5
9a6f2e3fea1b39bbb7343445db1167f7
License
GNU GPL v3 or later

The question each model received

Every model saw the same one-line instruction:

Correct the Korean sentence. Reply with the corrected sentence only.

Hosted models ran through OpenRouter. Offline models ran through LM Studio. Temporary failures were retried up to three times. Models kept their normal response settings; I did not force the same amount of internal reasoning on every model.

Runner
kolla-benchmark 0.1.0
Run dates
September 10–12, 2026

Scoring whether an answer was right

The scorer compares each answer with both human corrections for that sentence. If either comparison found a valid correction, it keeps the better one. That is the usual approach for sentences with more than one valid answer, though it gives models the benefit of the doubt.

The scorer used here is M2 MaxMatch (Dahlmeier & Ng, 2012). It is checked on every run against the reference Python m2scorer.

Useful changes, mistakes caughtPrecision, recall
How often a change was needed, how many mistakes were found, and one combined score that gives extra weight to avoiding bad advice.
Exact sentenceExact match
The output exactly matched one complete human correction.
Fully correctNo FP, no FN
No extra change was made and no needed change was missed.
No changesUnchanged
The answer repeated the learner’s sentence verbatim.
No answerFailed
The model did not return an answer after retries.

On cost and coverage

Hosted costs are the recorded service bill. They reflect exact prices paid on execution, not the monthly price of a consumer subscription.

Offline models have no hosted-service bill, but I did not measure hardware, electricity, or time, so they do not get any cost attribution.

A test is complete when every one of the 1,418 sentences has either an answer or a recorded failure.

This list does not contain every model available, especially not for local runs. The selection for local models was arbitrary, based on what was available to me at the time of testing. The selectino of hosted models was based on the current available services and their popularity.

What the numbers mean

Before comparing models, it helps to know what counts as a good correction. Here is the short version, in plain language.

Sometimes two corrections are both right

Two people corrected every sentence independently. When their answers differed, the corpus accepts either valid reading—because Korean often has more than one natural way to repair a sentence.

What the learner wrote

비행기 음식이 안 막였습니다.

One valid reading: “didn’t eat”

비행기 음식을먹었습니다.

Another valid reading: “didn’t suit me”

비행기 음식이 안 맞았습니다.

One learner sentence, with the two valid repairs used for scoring.

Doing nothing is not a zero

574 of the 1,418 sentences were already correct, so repeating them unchanged is the right answer. That means a model that changes nothing gets 40.5% of the sentences fully right—but it fixes none of the mistakes.

Good feedback catches mistakes without inventing them

Useful changesPrecision
When the model changes something, how often that change was actually needed. If this number is low, the model is more likely to rewrite Korean that was already fine.
Mistakes caughtRecall
Of the mistakes in the sentence, how many the model also found. If this number is low, it is more likely to let errors through.
Overall qualityF0.5 score
Our main score, which rewards both and gives a little extra weight to avoiding bad advice. The technical name is F0.5.

Each model was run once, so nearby results are effectively tied. Use the table to compare broad differences, not to declare a winner by a point or two.

Check the sources and raw data

Feel free to reproduce this benchmark on your own. The code and dataset are open source, and the raw runs results are available for download for reevaluation.

The raw reports contain each sentence, each answer, and the detailed scores.

Read the Personal Version or Try the Result

A benchmark can point you in the right direction. Try Elephant’s sentence analysis on Korean or read my personal telling of how and why I ran this benchmark.

Read the Blog Post Try sentence analysis