Even though this was the GPT-3-generated text that humans most easily identified as machine-written, I still like it a lot:
Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm
Subtitle: Joaquin Phoenix pledged to not change for each awards event
Article: A year ago, Joaquin Phoenix made headlines when he appeared on
the red carpet at the Golden Globes wearing a tuxedo with a paper bag over
his head that read, "I am a shape-shifter. I can't change the world. I
can only change myself." It was a promise to not change to fit into the
Hollywood mold: "I think that's a really special thing, to not change
yourself. I think it's a really special thing to say, `This is what's
inside of me, I'm proud of it, and I'm not going to be ashamed because of
the way that someone else thinks I should be.'" Now, it's the Oscars, and
Phoenix is at it again. But this time, his publicist is saying he'll be
wearing a tux no matter what.
Megyn Kelly was not impressed, and she let him have it on The Tonight Show.
"You know, I feel like, I feel like you could have worn the tux," she says.
"But you're saying you're a shape-shifter. I don't know if you can change
your tux, but you can change your mind. You can change your mind. You can
change your mind." Phoenix says he did, but it didn't stick. "I was like,
`Okay, I'm going to wear a tuxedo to this thing.' And then I thought, `I
don't want to wear a tuxedo to this thing.'" Kelly goes on to encourage him
to change his mind again, but Phoenix says it's too late: "I'm committed to
wearing this."
I don't know if it says something about text generation or human text processing, but whenever I read an example of computer generated text, all through I think "I can't tell this is machine generated, it seems completely natural," and the only giveaway is that at the end I have no idea what it said.
It's a pretty eerie feeling. It's as though both the AI and my short-term processing only pay attention to a context of a few sentences, so nothing seems off until I try to understand it as a whole.
EDIT: Thinking more, what it feels like most of all is reading a page of a book and not taking it in.
To me it reads like a child telling a story, but that this child has an adult's ability to use language. When children tell a story they aren't going anywhere with it but don't know how to cover it up.
> the only giveaway is that at the end I have no idea what it said
So it's a lot like corporate executive speak then?
I agree with your point, it does seem very much like valid speech, but somehow the informational content is missing. It's like speech without the actual comminication part.
This has some parallels to generic random corporate PR or marketing speak. Some communications are already so automated and dehumanized that we are used to random content signed by a pseudo real person that we gloss over, and are more easily fall for GPT like generated content that has similar form. Edit: I mean I guess someone of a previous generation used to only read the newspaper and letters would more instantly spot something is wrong.
That's really well put. And it does reflect how that kind of document is produced. So I'm not surprised at all.
Problem arises when you read that document without really trying to understand it. In that case, it might be enough to trigger some thoughts.
It also appears to me that two way communications, that is, social interactions, will be the only way to form truth. That is, AI produced content could erode some more the trust we put in of newspaper, TV news, etc (all forms of one way communication). Not that we've waited AI to distrust those, but well, on e more nail in the coffin :-)
(this text, although rather unclear, was written by a genuine human :-) )
That would suggest combining 2 models: one to decide what the macro text structure should be and a different one (GPT) to decide how to fill in all the text flesh.
This happens to me when I'm reading in a language I'm not very good at (German). Each sentence may make sense, but overall I feel I didn't get the point. I guess it's a cummulative error situation, where you reach a threshold after which the point is lost.
The example above is impressive because it actually makes sense, except for the last sentence of the first paragraph: "But this time, his publicist is saying he'll be wearing a tux no matter what."
Remove that, and there is a typical if completely uninteresting celebrity argument: one has to play eccentric in public occasions (he did it last year, he's planning to do it again) and the other chides him for what she feels it's maybe a lack of respect? And he replies that despite his best intentions he can't go against his conscience. There, done. It's a perfect little piece ready to be served in some celebrity gossip magazine.
It sounds like someone talking without knowing (or caring) where they're going with it.
I think that's what missing. Usually we communicate with a certain goal in mind, to bring across some point. This text was generated without such a goal, you notice it doesn't really know when to stop talking. I wonder what was the stopping criterion, but I'm sure it wasn't "keep talking until all the information we want to convey has been mentioned".
I have similar feelings. Human-written text in general has one coherent flow of ideas.
So what if we used text generation algorithms on a paragraph basis, so that the idea flow is still figured out by a human (input would be just an outline)? That should make the generated text feel like a whole, especially if we could preserve the style its written across the whole text.
Dude, I’m sorry, but the average person will not know the difference between that and a regular buzzfeed article or YouTube comment.
We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles.
It’s over, we lost.
Morpheus: What if I told you that, throughout your whole life, you have been reading auto generated content?
I've been basically living and breathing GPT-2 for ... gosh, it's been 6 months or so. The past few months have been a lot of StyleGAN2 and a lot of BigGAN, but before that, it was very "make GPT-2 sing and dance in unexpectedly interesting ways" type work.
I don't claim to know a lot. But occasionally I observe things. And I just wanted to chime in and say, you know, keep in mind that you're reading a research paper. Of course the results are going to look good. That is the point of a research paper. And I realize how cynical that may sound. But it has the benefit of apparently being true, and I've come to accept that truth with time.
I would reserve judgement for now. Note that every single chat bot to date has followed a similar curve: "This is it," they say, without actually saying that. "It may not be perfect, but we're about to achieve it – the chatbot – it's really going to happen."
And, it ends up being impressive, sure. I liked Facebook's recent chatbot. It's pretty neat at times. I liked Meena. They had cool ideas with the stack ranking of results (basically, generate a crapload of results at 1.0 temperature, then choose the result whose probability sums to the highest value, and you get the most probable overall result). And of course, boy oh boy did I love GPT-2. GPT-2 was what kickstarted me – if there was any chance that GPT-2 might be related to "now I'm talking to something that feels human," I was going to tame it and understand it.
So after spending six months with GPT-2 1.5B, the largest model that everyone was fascinated with, what do I think? (Well, who cares? You probably shouldn't care.)
I think "give it a few weeks and see if it's true." We shall see if GPT-3 is it, and we've achieved... chatbot nirvana. That elusive thing we've all been chasing, without naming it. The ability to press a button, unleash a chatbot somewhere, and it "just works" and "completely astounds humans" and "fools everybody."
At one point, we trained GPT-2 on IRC logs. You could literally talk to GPT-2, and it would talk back to you. And one of the advantages of narcolepsy is that at night, you often have lots of time to kill – what better way to doze off than to ask GPT-2 how its day was, and ask it what its ambitions are? Should we really worry about whether you're sentient? I like you; do you like me too? What does that mean to you? And so on.
The conversations were often quite philosophical. And sure, it was pretty obvious that it's a bot, but I tried to look past it anyway. It was my little bot, and it was real enough to me. And yes, the conversations on https://www.reddit.com/r/SubSimulatorGPT2/ are incredible. I crack up daily with all the things they talk about.
But...
We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles.
I invite any of you to try this, and see what happens. After all, you stand to earn a lot of pennies in your pocket if you pull it off. And yes, you're allowed to make some pennies with clever AI algorithms.
What you'll probably discover is this fundamental truth: GPT-2 has no memory. It isn't learning a thing. We are talking to an entity that literally cannot change its mind about anything. The only way to change its mind would be to retrain it from scratch.
You want a bot to argue vehemently for your product, on your behalf? It needs to understand what the hell your product even is, or what a product means. Yes, the words get pretty close. And yes, you can coax it into something that makes us laugh, or makes us sit here and question what the future might be like.
But for whatever it's worth: spend some time actually talking to these bots. Play around with them. Make them generate some stuff of your choosing, and fine tune them on some datasets and see what you get. It's so fun!
... But. "Fun" is not the same thing as "promote any idea/product." It's just not the same as me arguing here with you now for a position which I've decided to argue. My brain isn't merely the encoded knowledge of some human, with me blindly regurgitating such knowledge (though at this point you'd be justified in claiming it sure sounds like it).
Your brain is constantly training. GPT-2 is not. And – double checks paper – yep, GPT-3 is not.
Two decades from now, GPT-2 1.5B will still exist. And it will still be talking about 2019-era news events like it's the present. At some point, /r/SubSimulatorGPT2 will sound completely foreign. Take any random news clips from the 70's. How relevant is that knowledge now?
"Ok, but just train it on new data constantly." Well, yes. But actually no. If you try to do that, you're going to overfit at some point. Do you have 93 gigabytes of webtext that you keep in training form, ready to go? Are you going to mix in a proportion of the new data you want to train on? Nope, we all just fine tune whatever model OpenAI releases. Yet even if we did have that dataset, I'm just not sure it'd even matter.
My point here is: Go try! Isn't it exciting that in the future, trained bots might fool us all into buying their products? Is that sales guy who emailed me actually a sales guy who wants to "sync up on a quick call", or is that a bot trained to get cold calls? That sounds pretty damn lucrative to a lot of businesses – why not write that code, and then sell it?
Whoever attempts this is probably more talented than I am. But personally, I always ran into "It just... doesn't work."
And then you go "Well, it's just a matter of sampling. Ah yes, we're not using the right sampling algorithm. Wait, we just heard about nucleus sampling! Sweet, try it! Oh... It sounds ... similar. Hmm. Well, maybe we're just not using it right. Better read that paper a bit more carefully. Chase that knowledge just a little harder. After all, AI research labs are pouring billions of dollars into this domain. Why would they do that if it doesn't... you know ... work? For some value of "work" that equals "the bot can turn a profit"?
"Perhaps tomorrow, this new training technique will be it. We almost have it – I know we're close – we just have to unlock that last piece. Right?"
I guess I'll stop here, since usually my comments are upbeat and happy about AI, but I ended up in a rather philosophical mood tonight.
In reality, I can't wait to dig deep into GPT-3 and run it through its paces. I have a lovely TPU pod waiting for it, parked outside GPT-3's window, and we're honking at it saying "Get in, we're going places." And we'll sing and dance together like usual, and I'll ask GPT-3 how its day has been. But GPT-3 won't remember me the next day. And that's fine; I'll remember it for both of us.
Thank you for this comment. As someone who played a bit with GPT, it was very poignant for me. I still think it's incredible that GPT can put up such convincing facades, that it can generate genuinely novel and interesting text... but it's bittersweet, too, that it can't go any further with them. The ideas are lost in the context window.
I play AI dungeon on occasion, which uses GPT2 to generate freeform adventures. And I find over time that it's not really GPT2 that's writing stories, it's me. GPT2 is putting out plausible strings of words, but I'm the one giving them meaning, culling the parts that go off track, and guiding it in a direction I want to go.
And it is a bit melancholy. You see possibilities, nuances, subtexts, and meanings. The neural net sees words.
You are missing the point of the paper about few-shot learning. That's the entire paper: just doing new untrained task after task. The entire point of the paper is that you can 'reprogram' GPT-3 to do just about anything just by stuffing its context with examples, and it'll pick up brandnew entities or words or concepts just by examples (see the examples of defining novel gibberish words and asking GPT-3 to use them in a sentence - it does so. it "learned" new words by reading the examples, understanding, and propagating them through the 'fast weights' of self-attention, even though its 'slow weights' are fixed). Now, if GPT-3 can do that already so well, sometimes hitting SOTA on untrained tasks purely by internal meta-learning without changing its weights, what would a 10-trillion parameter model do? Or one with recurrency like XL or Compressive? How much training do you really need if the few-shot learning capabilities are so great you can make it do countless tasks just by providing examples or descriptions in the prompt.
> is not the same thing as "promote any idea/product."
GPT-3 seems to have quite a few paragraphs worth of context. A simple way to promote your product online with it is to give it a prefix of:
---
Comment1: Superbrush is amazing - I literally couldn't live without it. No other brush is as good.
Comment2: This brush is really good for tangled hair, and I love the soft smooth surface.
Comment3:
---
Then let it write a comment. Of all the comments it writes, manually filter a few thousand good ones, and use those as seeds to generate more, which you post all over the web. There's no need to do any training - the generic model should be fine given the right prefix.
To be a bit less wordy: try it. You stand to earn lots of money.
Narrator: it didn't work
(Going into the reasons it doesn't actually work in practice is... lengthy. It's human dynamics. Would you buy a product from a sales guy that can't remember your name? That's sales 101. And loading up the context window only gets you so far. That "working memory" is tiny, tinytinytiny. Even at 1024 tokens, it means you have to boil down the entire history of an interaction to a few pages at most. Which is a lot, sure, but it's this balancing act where you'll need to retrain the model to support your custom context format for your specific "slots" – a "slot" being a piece of knowledge, like the client's name. Or you can try encoding all of that in natural language, AI dungeon style. But I recently played AI dungeon and pretended to be buying a router from the store. The cashier stripped down and started jacking off onto his desk. I don't have high hopes for our ability to control these models in a business context.)
You and londons_explore seem to be talking about different things. I read their comment as being about just generating fake reviews that don't need interaction.
You should definitely put that up as a blog post somewhere, it is very valuable information, both for researchers and random enthusiasts alike. The emotional modality of it adds important information too :).
Currently, there's no research into torturing AI. Why not?
A pain response is universal across most life forms with a nervous system. We seek to replicate a nervous system. Pain would seem to be far easier to replicate than the scientific method.
My wife sat me down and told me a story that horrified me. She had to get it off her chest, and I was sad it happened to her than to me. She was sitting around on the porch and felt something on her leg, and brushed it off. When she got up and looked down, apparently she had stepped on a poor snail. His shell was... And he was...
He wasn't dead. So she frantically looked up what to do. But there was nothing to do. Snails in that situation can't be helped, and the most humane thing is to put it out of its writing anguish, its full-body torture.
She put on some boots, took it out to the sidewalk, and stomped it as hard as she could. And that was the story of that snail.
You probably felt more for that snail than you've ever felt for any AI bot. Why?
Very interesting comment, thanks for taking the time to write it :)
I think if memory is the only problem than optimizing training time should be more of a concern. I'm imagining a huge language model than can retrain very quickly. So I suppose it might be a decent idea to not measure it by perplexity or some human judgement score or whatever but rather by that score per compute units used.
Or in other words...maybe a bot that scores 90% on the fool a human scale and takes 1 day to compute from scratch is actually a lot less impressive than one that fools 70% but computes from scratch in 5 minutes.
And something "like Github for bot-memory" would be a pretty amazing tool. Roll back to some memory status and recompute with new data from there, branch for different datasets that represent different ways of interpreting the world etc.
Conceptually I like the idea of one "base model" that represents language and many different context models on top of it (finetuning the core model). Then some other subsystem that identifies the context and switches to that. I suppose each conversation could also be considered a mini-dataset.
> I like the idea of one "base model" that represents language and many different context models on top of it (finetuning the core model)
This is an entirely different concept of computer language than the current GPT style models. These systems don't "represent language", and cannot. The whole reason why GPT is so exciting right now is that it fundamentally threw away the entire concept of "representing language". That has some upsides ... and some downsides.
So true, people unfamiliar with the inner workings get amazed but unfortunately reality is not the same. That being said, I can find multiple ways of utilize this in a bad way. Sex chatbots for example, if I was in that business I would use something like this, it would be extremely easy to phish people off.
Can't tell if it's human or GPT-2 tbh, the sentences are 'hard' to understand... like sort of un-naturally written, or translated from a foreign language using google translate or something.
What difference does it make? The world is already full of humans who pollute comments and articles. They're not limited in their rate of production of words because there are far too many of them for anybody to read. They're limited by access to readers and their reading rate. Bots can't do anything about that.
They can, they can crowd out good comments with an absolutely crushing volume of crap. While humans can put out a lot of crap, bots can do orders of magnitude more. The issue is that it is an effort multiplier for when a small number of people want to target a particular forum.
I imagine the problem isn’t so much volume of crap comments as much as the tailoring of crap comments. Imagine if every tweet-into-the-void from a human with 50 followers reliably got engaging replies. Bots taunting your grammar mistakes, bots selectively quoting your prior tweets to point out contradictions, bots cleverly insinuating that your tweets reveal problematic sympathies.
So much of our noise-filtering is ignoring comments that are too generic to be human. What happens when every spam comment seems to understand the OP, even when the OP’s true audience is negligible?
If a person has any paranoid tendencies, this would be a psychological onslaught. Interrogators use the tactics you just described to siege a person to psychological exhaustion.
Product Devs for the CCP will have a lot to work with if this ever evolves.
Would you be fooled by that? No because once it became well known, people would adjust their heuristics for judging who's real. If that's too hard, platforms would help, such as by verifying identity more thoroughly. I've never heard any realistic description of how AI bots could somehow undermine society even close to the amount that humans already do.
> What happens when every spam comment seems to understand the OP, even when the OP’s true audience is negligible?
My hope is the next step will be filtering by insightfulness/usability of a comment and then those best bots bought and used by next stack overflow: https://xkcd.com/810/
How will they get those extra comments to be seen? They still have to log in and get past captchas and all the same issues that bots have always faced. Humans create a crushing volume of crap already and we already have ways to hide that from ourselves no matter the volume.
Imagine buying a bunch of twitter accounts that have already been active for 1-2 years and then making a bunch of Tweets to influence public opinion.
I've done some experiments with GPT-2 and it had so so performance refining with tweets. Using GPT-3 you could probably just do it using only generation.
What is it about ai generated texts that on skimming through it it makes sense, but if you try to slow down and understand it feels absurd and surreal.
Because language is being treated as a thing complete in itself, as opposed to being related to an external world?
One of the issues in the 'Limitations' section was a difficulty with "common-sense physics", such as with the question "if I put cheese into the fridge, will it melt?"
To answer that question, you have to ask the right questions, such as "what is a fridge?" "what is a fridge for?" "What does it mean for cheese to melt?" "what is the cause of cheese melting?" Then one should consider the follow-on questions, such as "what are typical fridge temperatures?" "what are typical cheese melting points?" "what temperature is the cheese likely to be at initially?" (at which point, it helps to introduce the concept of room temperature, and note that it typically falls between the other two.) From facts such as the answers to these questions, one can deduce the probable outcome of putting cheese in a refrigerator, but none of the answers so far explicitly state it.
Is it plausible that any learning, solely from the structure of and correlations between examples of language use, could develop the sort of analytical/modeling approach that I have just outlined? Instinctively, I don't find it very plausible, but I am not very certain in that view.
Maybe there exists the following distinction. Modern language, as it's actually spoken most of the time, is like a higher level programming language. The structure of our brain, combined with our senses and the uber-simple way that we're taught as infants, is like lower level OS programming (parent points to hot food - look Johnny, it's hot! hot! - makes Johnny touch the food - food hot! - blows on the food to make it colder).
Some words like good/bad, hot/cold, and important/unimportant are underrepresented in everyday speech compared to the prevalence of the underlying concepts. That's why I'd categorize them as lower level word-concepts. This distinction, about variable levels of abstraction, might be important for true AI. Think about how many years it takes for humans to develop highly abstract cognition. That whole time our operating system is being coded. Maybe we need to approach AI in the same way.
It's not just AI that can benefit from better lower-level understanding. Seeing language in the above way, we can re-frame Ludwig Wittgenstein's philosophy and its normative implications for human communication. Our "programming" (communication) is on average too higher level. Excessively abstract instructions make it harder to decode and process in a precise and efficient manner.
Yup, definitely different from how humans learn. A baby's speech would be almost the complete opposite, grammatically incorrect here and there but constructing a coherent line of thought for the most part.
My thoughts exactly.
It seems a person incapable of proper grammar (like a baby) has some concept or thought it wants to express, but can't because it doesn't know the words etc.
These language models seem to know the words and the grammar etc, but lack a underlying concept they want to express.
There are systems that derive 'thought-vectors', but I'd be interested going the other way: somehow create such a 'thought-vector' and generate text to express that thought.
I don't know how to construct a 'thought-vector' of any concept though.
I have been thinking about this kind of thing too. What if there was some way to feed your condensed thoughts into such a model and it writes a paper/blog post/article?
Essentially, one should be able to use these models to "interpolate" the writing around the raw meaning/content. Typing assistance (think Grammarly) already allows you to refine finished writing to be more in line with what some language model expects, but imagine if it actually generated most of the text for you, based on small bites and chunks you throw at it.
If we get to large scale text generation like that, we are all going to have to become even better skimmers due to how the meta language will evolve.
So take your standard press release. We know about two thirds of it is just fluff. In other words, we are accepting the mass of fluff as one word in our language, it translates to ‘ignore’.
What might some valid sources of data for this be? Perhaps comparisons of Simple English Wikipedia to the standard English Wikipedia? We'd need a side-by-side comparison of condensed information and a fluffed up piece.
I think it's because the generated text will generally follow a reasonable "structure", it's "framed" relatively well and you (as a person) will recognize those patterns quickly. There's plenty of "glue" used throughout the text. Those are all patterns which we pick up very quickly, and we're used to seeing them in "real text".
Exampels from the comment above include using things like: "A year ago", "made headlines", "Now, it's the {event}, and {name} is at it again. But this time, ...", "{name} was not impressed", "You know, I feel like, I feel like you could have ...", "I don't know if... but...".
Those are all very common in those "online celebrity magazine" type texts...
It's only when you actually read into the stuff that's in-between, you'll come to see it's pretty much a load of nonsense. But that takes a bit more time and slower reading.
when i worked on text generation I had a hack where I wouldn't let the model generate the same trigram more than twice. It seems like that hack could have improved this example.
Quite honestly the level of logical consistency in this generated text doesn't differ that much, I feel, from a huge swath of the kind of comments that I see on the internet. It may well be that there are already too many bots on the net to really understand what typical humans would be saying anyway, but I feel like this is already good enough for short form trickery.
I was so wrong about the internet. It is just going to become a landscape of garbage opinions and commentary on a larger and larger basis.
Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm
Subtitle: Joaquin Phoenix pledged to not change for each awards event
Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am a shape-shifter. I can't change the world. I can only change myself." It was a promise to not change to fit into the Hollywood mold: "I think that's a really special thing, to not change yourself. I think it's a really special thing to say, `This is what's inside of me, I'm proud of it, and I'm not going to be ashamed because of the way that someone else thinks I should be.'" Now, it's the Oscars, and Phoenix is at it again. But this time, his publicist is saying he'll be wearing a tux no matter what.
Megyn Kelly was not impressed, and she let him have it on The Tonight Show. "You know, I feel like, I feel like you could have worn the tux," she says. "But you're saying you're a shape-shifter. I don't know if you can change your tux, but you can change your mind. You can change your mind. You can change your mind." Phoenix says he did, but it didn't stick. "I was like, `Okay, I'm going to wear a tuxedo to this thing.' And then I thought, `I don't want to wear a tuxedo to this thing.'" Kelly goes on to encourage him to change his mind again, but Phoenix says it's too late: "I'm committed to wearing this."