UTF-8000: Unlimited UTF-8
23 points - today at 5:15 AM
SourceComments
deleted today at 7:34 AM
Sharlin today at 7:33 AM
UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
sph today at 7:11 AM
> UTF-8000 is in no way endorsed by or representative of the Unicode Consortium.
Not until they decide to expand the emoji range, allocate space for all past and future fictional languages, as well as birdsong and dog barks.
Someone at the consortium is rubbing their hands with glee with all the newfound space.
But honestly, cool hack! If you invent a method to encode large numbers into bytes, why limit yourself to 24-bit numbers?
achille today at 7:18 AM
> Ken Thompson: "...i really dont think it is useful. it is like replacing ipv6 with ipv50"
mrlonglong today at 7:26 AM
I love it.
Some day we'll need this when we finally realise we are not alone in the universe. Alien glyphs ftw.
Dwedit today at 7:12 AM
FF bytes are an easy way to identify an invalid UTF-8 file. This idea doesn't have that property.
flohofwoe today at 7:20 AM
Phew, and I was worried that we'd be running out of UNICODE space for new emojis ;)
Grimeton today at 7:29 AM
More like WTF-8.
snvzz today at 7:19 AM
No project is ever safe from complicators.
This is why we need the KISS enforcers.