Hacker News

rbanffy
When Compilers Disagree About UTF‑8 nemanjatrifunovic.substack.com

kstenerud5 days ago

You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop.

Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities.

vlovich12319 minutes ago

FWIW I believe the second optimization is at odds with the first. It’s really hard to do this kind of conditional historical check in SIMD.

duskwuff2 hours ago

> So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints.

Depends on what kind of text you're processing. Many languages use the ASCII range for spaces/newlines, digits, and punctuation.

mort962 hours ago

That, plus markup, be it something XML/HTML-like or something Markdown-like.

zX41ZdbWan hour ago

This is one of the optimizations from ClickHouse - detect ASCII and go a fast path: https://github.com/ClickHouse/ClickHouse/blob/6cde32de1a2463...

hn-front (c) 2024 voximity
source