Rendered at 10:47:28 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
MiroslavPokorny 10 hours ago [-]
Title is misleading, the compilers do not disagree, running the compiled code produces the same output. The difference is one binary(instructions) are slightly different...
kstenerud 5 days ago [-]
You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop.
Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities.
duskwuff 14 hours ago [-]
> So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints.
Depends on what kind of text you're processing. Many languages use the ASCII range for spaces/newlines, digits, and punctuation.
mort96 14 hours ago [-]
That, plus markup, be it something XML/HTML-like or something Markdown-like.
vlovich123 12 hours ago [-]
FWIW I believe the second optimization is at odds with the first. It’s really hard to do this kind of conditional historical check in SIMD.
Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities.
Depends on what kind of text you're processing. Many languages use the ASCII range for spaces/newlines, digits, and punctuation.