Skip to content

kz

Rust · CLI

A parallel replacement for GNU wc. It's 10 to 68× faster on real text, and its counts match wc exactly in every mode.

GitHub crates.io/kz-cli kz-cli 0.2.0 (c54495d)

Two numbers I had to take back

I used to publish two numbers for kz that were wrong. Both came from benchmarking on generated text without checking the output. I only caught them when I measured again on real books.

-L measured bytes instead of display width

café naïve résumé is 17 characters wide. kz said 21 because it counted bytes. The old 88× longest-line number came from this. kz was doing a quick byte scan while wc worked out the width of every character, so the two weren't doing the same job.

Encoding detection read the whole file

A single accented character in a 97 MB file made word count go from 21 ms to 5.36 s. With generated ASCII the fast path always ran, so it looked like 21×. With real prose it almost never ran, and kz was 4.6× slower than wc.

Both are fixed in 0.2.0, and every number on this page was measured after the fix.

Now the benchmark runner won't time anything unless kz and wc give the same counts. On top of that, kz has been checked against wc about 2,000 times, plus fuzzing.

10 to 68× on real text

-m, 100 MB of real prose. This split only holds for -m.

12.1×

algorithmic

Measured on one thread. wc -m decodes one character at a time with mbrtowc. kz validates UTF-8 over the raw bytes in one pass, which the compiler can vectorize. You get this part on any machine.

6.0×

parallelism

What 16 threads add on top of that. This part comes from the hardware. The chip has 8 physical cores, and they do most of the work here. SMT adds a little.

68 to 72×

combined

On -m over 100 MB of prose. Three runs all came in at 3.3 ms ± 0.1. One worked out to 68.4× and another to 72.6×, so I give it as a range.

  • -m, 100 MB real prose 68.4×
  • -L, 1 GB real prose 16.2×
  • -w, 1 GB real prose 14.7×
  • default, 1 GB real prose 13.1×
  • -w, 100 MB real prose 13.0×
  • 141 real files, 100 MB 11.6×
  • -m, 6.7 MB CJK 10.6×
  • -w, stdin 100 MB 7.1×
  • -l, --files0-from 2972 files 4.6×
  • -l, 1 GB real prose 2.8×
  • -l, recursive tree walk 1.4×

speedup vs GNU wc · linear scale to 72×

With LC_ALL=C, wc -m counts bytes instead of characters. It finishes in 501 µs but reports 104,329,273 instead of the correct 102,891,699. kz counts characters in every locale, so the 68× above compares the same work.

The split per mode

That split is only true for -m. Here's how the other modes do on one thread and on sixteen.

mode wc kz, 1 thread kz, 16 threads algorithmic parallel
-m 232.4 ms 19.2 ms 3.2 ms 12.1× 6.0×
-w 229.5 ms 163.0 ms 17.5 ms 1.4× 9.3×
-L 231.1 ms 130.0 ms 16.6 ms 1.8× 7.8×

Only -m gets a big win from the code itself, because wc -m takes a slow path that kz avoids. wc's word counting and line length code is already fast. On one core kz -w is 1.4× faster and kz -L 1.8×, so most of their speedup comes from using more cores.

Past a point, more threads stop helping. This is -m going from one thread to sixteen.

  • 1 thread · 19.2 ms ± 0.4 12.1× first row
  • 2 threads · 10.5 ms ± 0.3 22.1× 1.83×
  • 4 threads · 5.8 ms ± 0.1 40.1× 1.81×
  • 8 threads · 3.5 ms ± 0.1 66.4× 1.66×
  • 16 threads · 3.2 ms ± 0.2 72.6× 1.09×

-m, 100 MB real prose · speedup vs wc · right column is the step over the previous row

Going from 8 to 16 threads barely helps, because the 7800X3D has 8 physical cores and SMT only adds about 9%. On a four-core machine, expect around 40× on -m instead of 68×.

Limiting Rayon to four threads on an eight-core chip isn't the same as running on a four-core machine. The limited run still gets all 32 MB of L3 cache and all the memory bandwidth, so treat these numbers as the best case for four-core hardware.

Two cases where wc wins

-l, stdin 100 MB

wc ~1.2× faster

wc
8.0 ms ± 0.1
kz
9.2 ms ± 0.2

You can't memory-map a pipe, so the best kz can do here is match wc.

startup, 4 KB file

wc ~1.2× faster

wc
466.7 µs ± 52.1
kz
563.0 µs ± 59.9

This is process startup. Running kz --version alone takes longer than wc takes to count the whole file.

The startup row is under a millisecond and its spread is about 11% of the mean, which is mostly startup noise. That's why both margins are rounded to about 1.2×. wc is faster on files under about 100 KB and on piped input.

How it works

kz is faster for two different reasons, and which one matters depends on the mode.

The first only applies to counting characters with -m. GNU wc does this with mbrtowc, decoding one character at a time, and each step has to wait for the one before it. kz validates UTF-8 over a plain range of bytes instead, which the compiler can vectorize. That's where the 12.1× on one thread comes from. wc's word counting and line length code doesn't have this problem, so the other modes don't get the same boost.

The second is parallelism, and every other mode depends on it. Files are memory-mapped and split into chunks, and Rayon counts the chunks in parallel. If a multi-byte character ends up split across two chunks, it still gets counted once. In these modes kz does the same work as wc, just on more cores. On one thread -w is only 1.4× faster and -L 1.8×.

That's also why kz loses on small inputs. You can't memory-map a pipe, and a 4 KB file is finished before the extra threads pay off. Below about 100 KB, wc is faster.

Full results

workload wc (UTF-8) wc (LC_ALL=C) kz result
-m, 100 MB real prose 232.4 ms ± 4.2 501.0 µs ± 105.4 3.4 ms ± 0.1 68.4×
-L, 1 GB real prose 2.309 s ± 0.008 1.770 s ± 0.021 142.5 ms ± 1.4 16.2×
-w, 1 GB real prose 2.280 s ± 0.027 1.780 s ± 0.030 154.6 ms ± 1.4 14.7×
default, 1 GB real prose 2.270 s ± 0.028 1.781 s ± 0.027 173.6 ms ± 1.5 13.1×
-w, 100 MB real prose 229.5 ms ± 3.3 180.6 ms ± 4.4 17.7 ms ± 0.9 13.0×
141 real files, 100 MB 232.9 ms ± 0.8 not measured 20.0 ms ± 0.3 11.6×
-m, 6.7 MB CJK 22.2 ms ± 0.3 359.6 µs ± 79.2 2.1 ms ± 0.2 10.6×
-w, stdin 100 MB 232.5 ms ± 2.6 not measured 32.6 ms ± 0.5 7.1×
-l, --files0-from 2972 files 23.9 ms ± 0.6 not measured 5.2 ms ± 0.2 4.6×
-l, 1 GB real prose 53.0 ms ± 0.7 not measured 19.1 ms ± 0.4 2.8×
-l, recursive tree walk 34.4 ms ± 0.6 not measured 25.3 ms ± 0.7 1.4×
-l, stdin 100 MB 8.0 ms ± 0.1 not measured 9.2 ms ± 0.2 wc ~1.2× faster
startup, 4 KB file 466.7 µs ± 52.1 not measured 563.0 µs ± 59.9 wc ~1.2× faster

Mean ± standard deviation. A dash in the LC_ALL=C column means the workload doesn't decode characters (line counts, file lists, startup), so the locale can't change the result and I didn't measure it.

hardware
AMD Ryzen 7 7800X3D · 8C/16T · Linux 6.17.7-zen
harness
hyperfine, 3 warmup + 10 measured runs (50 + 500 for the sub-millisecond startup row), page cache warmed before each row, against GNU coreutils 9.11.
corpora
141 distinct Project Gutenberg texts concatenated (104 MB, 2,011,399 lines, 17,650,663 words, 2.15% non-ASCII); 6.7 MB of CJK text at 96.3% non-ASCII; and 2,972 .rs files vendored from kz's own Cargo.lock, so the code corpus is byte-identical on any machine.
scale
The 1 GB rows are the 100 MB corpus repeated ten times. Counting is a straight scan through the file, so repeating the text doesn't make kz look better. The 100 MB rows show how kz handles varied text, and the 1 GB rows show throughput on large files.

Reproduce: ./bench/prep_real.sh then ./bench/run_real.sh