I don't think it's meant as something that's useful per se because highlighting 5m lines of code at once is unusual, and doing it in a non-deterministic way probably isn't especially helpful, but to show what a browser is capable of it's awesome.
The use case isn't highlighting a huge amount of code, but local re-entrant highlighting.
You generally never need to highlight a whole file, only the part you're looking at. Which by definition must be fuzzy since you don't have the full source. And it's actually one of the nastier parts of writing a highlighter.
The lesson is that if you're going to be fuzzy you might as well be learned too and the results are pretty good. This is more useful than it appears at first glance.
> the part you see will not parse as a complete program so you need to guess
Presumably this could also be a useful trait when live highlighting of files when editing, as in-progress typing is likely to be unparseable sometimes.
Is that really the use case for the related packages mentioned on their website? highlightjs and prismjs are typically used to highlight code snippets on websites… which are at most a few hundred lines long. And are usually highlighted as a whole
No need to guess... For basic system highlighting, one can store the lexer context periodically (maybe per line) and invalidate it if edits are made before it. For more advanced stuff, you need a more complete parser anyway and do it asynchronously (LSP...).
That is a good optimization, but you still need to figure out the starting state at a given location, and for quite a lot of languages that is not solvable in the local case.
E.g. if you're looking at a page of Ruby, you can not in the general case know if it's inside or outside a quoted string, as the quote character can be any arbitrary character.
So unless you parse from the start of the file, you're left with fuzzy matching, and that can be made good enough the vast majority of time.
If anyone can get in contact with Shu, please let them know I'd really like to speak with them. I have a deterministic rendering / colorizing system in three dimensions instead of two for arbitrary UTF8 text written that is explicitly GPU bound and currently testing the last Rust + Mojo pair port, and I'd like to share some ideas. I have about 95 million glyphs rendering (monospace, atm) in a little under a second with full multi dimensional pagination and colorizing per-glyph with full addressing capabilities. There's something to be combined with these two mechanisms I'd like to try and explore.
27 KB is extraordinarily impressive. Here's a deterministic suite of common language- and format-specific PHP files an LLM wrote for me in 32 KB, for comparison:
I think what's most interesting is not that this works well for common languages, but it seems to be useful for languages that don't exist yet. If you want _some_ highlighting for your SQL variant, or your language that borrows many common idioms, or just sloppy code that's close enough to the target language, this could be useful.
This is a tiny LLM doing all the heavy-lifting. Any mention of the training process? I am obsessed with tiny LLMs and the do-one-thing-really-well approach that they seem to fit very well.
I think, both FF and chromium on Linux disable WebGPU by default. Of course, it depends on where exactly did you get it (what kind of default settings people that packaged it for your distro used).
Just look up "<browser name> enable webgpu" in you favorite search system or ask an LLM chatbot "How to enable WebGPU in <browser name>?", follow the instructions and it should work just fine.
Years ago, long before the current "AI"/LLM-craze, I've heard someone saying that in the future programming will have less fixed, deterministic algorithms and more statistical/ML algorithms, because even in the case when writing a deterministic algorithm is possible, it is sometimes easier to collect training data and train a tiny model, then to write and maintain large codebase that does the same thing, but deterministically.
I wonder, if mass adoption of LLM code-generation will accelerate this process or not. On the one hand, it is now easier to "write" (LLM-generate) code then ever, so time/effort savings are not there anymore (probably, I'm not sure). On the other hand, now everyone uses LLMs so, I think, people are less deterred by non-determinism and statistical nature of ML models.
Let's imagine that you are making a CLI tool. What does a CLI tool do? Well, it accepts reads arguments, reads stdin, writes stdout and make syscalls. So, all possible inputs and outputs are very well defined. What if in the future it will be easier (and maybe even more natural) to ask an LLM to "imagine" tons of possible inputs and correct outputs for a tool that you are making and then train a tiny model, without writing or generating any code?
"the future programming will have less fixed, deterministic algorithms and more statistical/ML algorithms"
I think the direction this will be going depends in part on how hardware prices will develop. If we continue with "hardware is cheap, don't think about it" like we did in the past decades I can see this happen. On the other hand, if we continue down the path we've recently taken, people who know when to choose which approach to their advantage will have a good and prosperous life.
> What if in the future it will be easier (and maybe even more natural) to ask an LLM to "imagine" tons of possible inputs and correct outputs for a tool that you are making and then train a tiny model, without writing or generating any code?
This seems to assume you're okay with whatever you're building being a black box that will break in the future and require you to re-train the model for every scenario that comes up. I can't think of a problem I've solved recently where that would pass the bar for me. Maybe for one-off problems like 'I have all this data and I want to classify it' where you could hand-classify say 5-10% of your dataset and train a model to deal with the rest of it?
A genuinely impressive and novel use of machine learning? In this economy?
Looks really impressive, although one issue I do see is consistency. For example, it seems to highlight `null` as a value, unless it is on the left side of an `===` expression, since it probably hasn't seen that in training.
Was it ever actually tested on languages not in the training set? Would be interesting to see if a model this small can actually generalize...
The bootstrap.css example shows a lot of very random highlighting around numbers (some are blue, others black) and rule names (some are purple, others black).
While the premise is neat (syntaxes and tokenizers are pretty similar, how much diversity can humans possibly come up with??), this does not inspire much confidence in the accuracy.
40 comments
[ 476 ms ] story [ 1590 ms ] threadYou generally never need to highlight a whole file, only the part you're looking at. Which by definition must be fuzzy since you don't have the full source. And it's actually one of the nastier parts of writing a highlighter.
The lesson is that if you're going to be fuzzy you might as well be learned too and the results are pretty good. This is more useful than it appears at first glance.
Presumably this could also be a useful trait when live highlighting of files when editing, as in-progress typing is likely to be unparseable sometimes.
E.g. if you're looking at a page of Ruby, you can not in the general case know if it's inside or outside a quoted string, as the quote character can be any arbitrary character.
So unless you parse from the start of the file, you're left with fuzzy matching, and that can be made good enough the vast majority of time.
Speak for yourself, I often accidentally open 50MB JSON files, crashing my text editor as it tries to figure out the syntax highlighting :)
Great news for a language whose last stable release was 20 years ago
https://repo.autonoma.ca/repo/treetrek/tree/HEAD/render/rule...
Is this a linux issue / support issue?
Just look up "<browser name> enable webgpu" in you favorite search system or ask an LLM chatbot "How to enable WebGPU in <browser name>?", follow the instructions and it should work just fine.
Yup, figured out how to enable and it works well now.
I wonder, if mass adoption of LLM code-generation will accelerate this process or not. On the one hand, it is now easier to "write" (LLM-generate) code then ever, so time/effort savings are not there anymore (probably, I'm not sure). On the other hand, now everyone uses LLMs so, I think, people are less deterred by non-determinism and statistical nature of ML models.
Let's imagine that you are making a CLI tool. What does a CLI tool do? Well, it accepts reads arguments, reads stdin, writes stdout and make syscalls. So, all possible inputs and outputs are very well defined. What if in the future it will be easier (and maybe even more natural) to ask an LLM to "imagine" tons of possible inputs and correct outputs for a tool that you are making and then train a tiny model, without writing or generating any code?
I think the direction this will be going depends in part on how hardware prices will develop. If we continue with "hardware is cheap, don't think about it" like we did in the past decades I can see this happen. On the other hand, if we continue down the path we've recently taken, people who know when to choose which approach to their advantage will have a good and prosperous life.
This seems to assume you're okay with whatever you're building being a black box that will break in the future and require you to re-train the model for every scenario that comes up. I can't think of a problem I've solved recently where that would pass the bar for me. Maybe for one-off problems like 'I have all this data and I want to classify it' where you could hand-classify say 5-10% of your dataset and train a model to deal with the rest of it?
Looks really impressive, although one issue I do see is consistency. For example, it seems to highlight `null` as a value, unless it is on the left side of an `===` expression, since it probably hasn't seen that in training.
Was it ever actually tested on languages not in the training set? Would be interesting to see if a model this small can actually generalize...
But what could I expect from another vibe coded project?
While the premise is neat (syntaxes and tokenizers are pretty similar, how much diversity can humans possibly come up with??), this does not inspire much confidence in the accuracy.