Not quite, UTF-8 encodes the length of the sequence in the first byte, instead of marking the end, this makes more sense because it is easier to detect broken input. 1110xxxx 10xxxxxx 0xxxxxxx obviously has a byte…
Not quite, UTF-8 encodes the length of the sequence in the first byte, instead of marking the end, this makes more sense because it is easier to detect broken input. 1110xxxx 10xxxxxx 0xxxxxxx obviously has a byte…