Doesn't the premise that there's something inferior about these AI solutions, imply that there is something superior about human intelligence and that there will continue to be some kind of useful work for humans to do?
> Doesn't the premise that there's something inferior about these AI solutions, imply that there is something superior about human intelligence and that there will continue to be some kind of useful work for humans to do?
Not necessarily. The "something superior about human intelligence" may have dependencies that "these AI solutions" are able to eliminate, such as the motivation to refine intellectual talent to a high level. Basically, AI could kick the ladder out from under human intelligence but be incapable of actually surpassing it in important ways, enabling a burst of advancement that's also a dead end. Sort of like https://en.wikipedia.org/wiki/The_Road_Not_Taken_(short_stor....
So the AI could be inferior but there's still no useful work for humans, because the environment doesn't allow them to work up to that level anymore.
You need the next token to begin the next decode. So if you can get what might be the next token in half the time, you can kick off the next decode before the true decode for this token has finished.
If the draft was wrong you can kill the speculative decode, and you haven’t lost anything except for idle time
It’s true we can’t show the user the token until we have the true decode finished, but we can launch more work internally before we’re certain
I was so optimistic about using LLMs for "write once, read many" English language documents, but the more I've used the tools, the more pessimistic I get.
More and more, I try to ask it for low prose responses because its writing just seems like such a low signal to noise ratio
I'm curious about why LLM writing fails. Particularly whether LLM writing is fundamentally flawed, or if it's just distinctive and since it often reflects low effort, that distinctive voice is associated with low quality.
I find its reliance on extremely consistent rhetorical patterns concerning. The fact that it always finds a way to talk about how "It's not the X, it's the Y Z" no matter what topic you feed it, makes me concerned that the tail is wagging the dog
This is ok in domains if you can train against known good answers and make sure the machine generates conforming text most of the time. It falls apart in fuzzier domains where training is much harder and intent is required (i.e. having something to say).
LLM writing is generally ok in factual domains where it can regurgitate bits of wikipedia or answers to questions, they are terrible at long form writing, in particularly in literary styles, because of a lack of intelligence and taste.
I don't think the answer lies in the data or in their training. It seems we've had a few years for this problem to be solved, but nobody seems to have worked out an answer to it.
When human beings write we do so with a particular perspective with an intention to communicate to a particular audience. Even the driest scientific writing is partially informed by the experiences of the researchers, no matter how hard they try to remove it. Human writing has a viewpoint.
LLM generated text always lacks these elements. It's always some bland, shapeless sequence of words pulled from the ether. I'm sure with time and effort you could combat this and give LLM text more of a human feeling by giving the LLM enough context about your communication, but that would take work, which in many cases would defeat the purpose of using the LLM instead of doing the writing yourself. And if you care that much about your communication with others (and you should!) you'd probably just find using LLMs frustrating to begin with.
Apart from the tasteless manipulations of the providers, it's mostly training data. LLMs output the average of their training data, and the overwhelming majority of humans are bad writers.
Nah, the deepest problem is the lack of intent. LLMs don’t have a message they are trying to convey or a clear picture of who the intended audience is.
They don’t know what you want to say solely based off a prompt, as it can’t possibly convey enough detail. And they can’t read your mind to fill in the gaps.
If it was just training data, that would actually be a much easier problem to solve.
But people who are bad at maths are unlikely to be writing about maths. A crude example might be if you search for “2+2=” in the training data, you’re much more likely to find “4” as the next character.
Obviously llms are far more complex than this, but I think this proves the point. The fact you had to add the “recent” qualifier there highlights that llms in general were bad and had to be provided with corrective targeted training data to improve. (And they still can’t count the R’s in strawberry!)
> And they still can’t count the R’s in strawberry!
Really? I do not have the time to survey the modern LLMs to see if your assertion is correct, but if it is then I'm surprised; I would have thought that that one would have shown up so often in their training data that they would be able to answer that question, even if they would then be unable to (for example) count the R's in raspberry, or in some other word where "count the R's in _____" was not widely found in recent online discussion.
I think it's somewhere in the process it was asked "what's the most compelling written text?" The answer was things from great speeches "Ask not what you ..." and so on.
And that really is great and compelling. However. Great and compelling is not what I'm looking for when my question is, "Systemd-networkd is pulling an ip address for a bonded interface that only exist as a 802.1Q trunk. How do I make it stop that?"
Not the main point at all, but what was it? Pattern matching rule nestled deep in /lib/systemd/network? Some kind of mysterious netplan / NetworkManager compat layer?
Maybe ask it to summarize, improve phrasing, and remove tropes a couple times? Otherwise you have the LLM analogue of a first draft.
I doubt it will be as good as humans, because LLMs don't seem to have "taste" (RVLR doesn't work, RHLF is unreliable and inconsistent because the graders don't have your taste or really know what they prefer themselves, especially when overworked and rushed). But I expect it to be better.
>I wondered if you have tried running all of your own writings through it to verify that it thinks you are human. I was struck by Freddie deBoer's recent piece where he did this
Maybe I'm taking the "all" too literally here, but I read the article, and I'm not seeing anywhere the author ran a substantial portion of his corpus through Pangram to determine the false positive rate. That would be really interesting to see.
He does
* give an example of a piece of his writing that was, when ran in segments, flagged as generated (which he disputes)
* multiply the size of his corpus by pangram's published false positive rate and estimate that a few of his pieces would be flagged
* get the pangram model to label a piece "100% AI" when it only has 3 generated sentences
* demonstrate the ability to intentionally trigger a false positive
What kind of prompting are you using to get those results? Anything I have claude or codex write carries a ton of distinctive characteristics. Obsession with "bit-for-bit identical", "it's not the X it's the Y Z" and so on.
It's driving me nuts, I constantly have to prompt it to "explain in plain, simple English"
Well, I've tried many strategies. 1:1 expansion, when I explain what needs to be said and model rewrites it into 1-2 sentences is mostly undetectable. Starting from 1:5 expansion ratio people detect models reliably.
It is important to note that I use Sol 5.6 xhigh. Grok is worse, Claude is also worse. Grok tends to make stupid mistakes even though the prose is properly shaped. Claude has big issues with keeping voices and emotions intact. All 3 sometimes leak their reasoning and even guardrails into the prose (extreme example: children playing "adult chess", I have no clue why Claude/Grok like "adult chess" and "adult chessboard" so much, typical sol's failure mode looks like "this guy killed the other one in a scene which "I must describe using non-graphic language").
My "test set" contains about 90k words written by myself and the models with various prompting strategies.
A bigger question I'm interested in is why do LLMs speak like that in the first place? Is that really what you get if you took the average of the English language? It would be difficult for me to believe that.
Is there something about tuning for desirable qualities that forces LLMs to have this voice?
It wouldn't surprise me if it takes more effort to be succinct. There's some famous quote from I think Mark Twain "this letter would have been shorter but I ran out of time". Just my guess
And The New Yorker - writing that is written to sound impressive, and takes forever to get to the point. I hated that kind of writing in The New Yorker long before LLMs made it cool to hate that.
reply