Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

GPT-3 demonstrates that a huge volume of what's written is mostly bullshit. This is very upsetting to some. See "The Digital Zeitgeist Ponders Our Obsolescence" in the linked article. What comes out of this system is better than most comments on political blogs, and sometimes better than the articles.

On what would this approach do badly? "How-to" material, I suspect. Trained on auto repair manuals, it could generate new, plausible, but useless, auto repair manuals. This gives us an insight into what's wrong. It lacks adequate ties to the real world.

This is the "common sense" problem I've discussed previously. Figuring out what's going to happen next in the real world is often not a problem in word space. It's a problem in a different kind of space. The shape of that space is a big unsolved problem in AI.



I think perhaps what's more upsetting is that GPT-3 flips the traditional notions of what machines are good at and what humans are good at on their respective heads.

GPT-3 seems to indicate there's a chance that "creative" domains such as poetry, literature, music, etc. will be taken over by AI (i.e. AIs will have superhuman performance) before "logical" domains such as logic, mathematics, and the sciences.

This means that it is becoming more and more conceivable to more and more people that sometime in the foreseeable future an AI will be better than any human along any dimension you choose to measure, even when it comes to the ability to elicit emotions and reactions in other humans.


I think you hit the nail on the head, with the salient point here being that in the near future "creative" things will be automated first (see Image GPT, Jukebox, etc. Google has 100 billion dollars cash and countless TPUs, best engineers, infra, etc - they could probably replicate results far better than each of these OpenAI projects within a few years). One of the things that got me into ML research was the notion that we could automate a lot of the hard work humans do every day (agriculture, cooking, desk jobs, etc) so that humans could do things that were uniquely theirs & interesting, that were human, that were beautiful... Unfortunately it turns out that classical music and waxing poetic are easily generative in an enjoyable way. In the most ironic fashion possible, it turns out that the very thing we do when we conduct ML research, what you call the "logical domain", is one of the only things that stays human-only in the foreseeable future.

GPT-3 and other projects seem to drive hype cycles in the tech community and convince people like Elon Musk that the AGI revolution is near. But I think recent progress is just another example of machine learning models being able to generalize on super large datasets, even if it's the biggest model so far. It's not clear to me that larger models will solve this in the limit; take the way GPT3 fails on addition past a certain number, and the fundamental inability for transformers to learn certain algorithms. It is certainly still possible for this type of large dataset, large model style of ML to make human life better in many ways - like Tesla is trying to do with self driving cars, or Covariant with automating Amazon-like jobs. But I think when it comes to tackling the hard problems of true intelligence, we're missing a dimension somewhere.


Disclaimer: I'm a composer

> Unfortunately it turns out that classical music and waxing poetic are easily generative in an enjoyable way

On the contrary, I would say that generating convincing and original classical is an incredibly hard (if not impossible) task. All the current music AI projects give results which may sound “good“ to a casual listener, but they sound horribly wrong to any educated listener. The reason is that AI can only imitate the surface, but completely misses to recognize/synthesize larger structures. This might be ok for some background noodling in a TV drama, but not for the concert stage.

Finally, we rarely perceive art works in isolation. We know and appreciate the fact that a certain work has been created by a certain person in a certain time.


The reality is likely neither here nor there - i.e. computing may have more to offer to the creative endeavor than creators would like to admit, but still leave an obvious gap which technologists might be loathe to admit.

It may be instructive to look at David Cope's [1] work (what he calls "recombinant music" [2]). Cope's been writing algorithms to compose in the styles of the masters (Mozart/Chopin/et al) for about 3 decades now, well before the recent surge in "AI". His techniques are much less sexy for the "deep learning" enthusiasts, and yet he managed to outrage an audience of connoisseurs who assembled to listen to a "lost Chopin piece" only to be told, after they shared their applause, that it was composed by a computer taught to mimic Chopin's style (the composition was performed by a musician). The response, in my opinion, also points to music as a social constructed experience and not purely attributable to the sound signal itself. i.e. if I give you a romantic background story for a lost composition of a master, you may be inclined to experience the piece in a more favorable light than if I told you it was generated by an algorithm (or the converse).

You're absolutely right that the musical output of the current crop of "AI" projects (especially the ones using deep learning / neural networks) are crappy to even a modestly trained listener .. or even a lay untrained listener for that matter. However, more involved modeling (such as Cope's) has produced some very compelling results decades ago, so it would be a mistake to assume that the current crop won't get close enough [3]. The fact that DL systems don't need to be instructed in the way Cope has had to encode his musical understanding is also something to be considered in the evaluation as well as in scoping their capabilities going forward.

[1]: https://en.wikipedia.org/wiki/David_Cope [2]: https://www.recombinantinc.com [3]: https://deepmind.com/blog/article/wavenet-generative-model-r... (see "Making Music" section and examples there)


I am also a computer musician, btw, so I am well aware of the creative potentials of algorithmic composition. ;-)

However, we have to make a clear distinction between creative and recreative methods. David Cope's work is impressive, but it focusses on the recreation of existing musical styles. This is interesting from a musicologist perspective, but not very interesting artistically.

I would certainly say that deep learning generates lots of interesting “material“ (like many other methods of algorithmic composition), but we still need a human being to curate, edit and assemble the material into a meaningful piece of art.

Finally, I think the current AI debate can be very fruitful for the arts. In a way, it raises similar questions as the concept of the “readymade“ and the pop art movement did in the 20th century.

Btw, I'm currently working on an opera which uses AI generated lyrics :-)


Humans also need other humans to curate their work. We are comparing AI not only to the best composers alive, but also to the best composers ever. Nobody remembers millions of failed musicians.

BTW - I'm curious, what do you think about birds songs? Are their songs interesting artistically? How do you think they were composed?


Oh, you're opening up a huge topic there. Actually, there have been philosophers who claimed that the beauty/sublimity of nature was ultimately superior to the sensations produces by the arts. You can find this reasoning in Kant's "Kritik der Urteilskraft", for example.

On the other hand, you have composers like John Cage (or more recently: Peter Ablinger) who claim that the act of listening itself can be/create art, blurring the borders between nature and art. There are conceptual pieces which only consist of listening instructions.

Finally, bird "songs" have been used as the source material for musical composition for centuries. You can find it in Beethoven, Mahler, Debussy, Stravinsky, etc. Olivier Messiaen even was a hobby ornithologist; he faithfully transcribed hundreds of bird songs and used them in his music (see for example his piano cycle "Catalogue d’oiseaux").

As for the question of who composed the actual bird songs, the answer probably depends on the theological background of the person you ask ;-)


I'm willing to go a little further with recombination given that a good part of a traditional musician's education consists of studying and re-performing "standards" be they jazz, western classical or Indian classical (which is my background). A simple example is how pretty much every hero-soundig film background music smells of Also Sprach Zarathustra to me. I do think that musician's stand as much on the shoulders of giants as scientists do .. but sometimes don't quite acknowledge that explicitly in their works.

I think this topic will keep reverting to the point you raise - "meaningful art". As long as the "meaning" is a construct in a human brain that we're looking for, we have little to say about AI and it's capabilities (like Joshua Bell's hardly-noticed playing of Bach classics at New York's subway station as opposed to when he's performing at a concert hall).

.. (edit) and I do think that active listening is itself a creative act.


> All the current music AI projects give results which may sound “good“ to a casual listener, but they sound horribly wrong to any educated listener

I think you're right, in that AI won't be able to create deeper themes and patterns, but I disagree with the above point: AI will take over the music industry because the vast, vast majority of people aren't educated listeners. The popularity of 6six9ine is a fantastic example.

To put it another way, I don't need another Terry Riley, Clint Mansell, or Meredith Monk, I just need something good enough to occupy some brainspace while I drive home after work; a move soundtrack just needs something sad, or exciting, or tension building. The AI can and will get there soon enough.


Even if it takes over the industry (I can actually imagine this happening), my original point still holds: the educated/experienced listener will notice and will care. For some people at least, music or art in general will always be an existential form of human expression, not some random exchangable consumer product.


All the current music AI projects give results which may sound “good“ to a casual listener, but they sound horribly wrong to any educated listener. The reason is that AI can only imitate the surface, but completely misses to recognize/synthesize larger structures.

Lack of "larger structures" is the key here. That's where GPT-1 was. Each sentence, in isolation, seemed to make sense, but after a few lines, it was clear the text wasn't going anywhere. By GPT-2, paragraphs seemed semi-reasonable, but multiple paragraphs didn't hold together. GPT-3 is able to keep it together for a few paragraphs, but probably not for a book chapter.

Music synthesis has the same scaling issue. Generators which imitate known patterns work for a few bars, but after a while you realize the music is going nowhere. The GPT results on text indicate that a scaleup may fix that problem.


This is the same argument people made against MP3 compression.

Lossy is bad. Humans will never stand for it.

Perfections will not stand for it. Pragmatists won’t notice.

This isn’t a bad thing. We need perfectionists to drag us across the “good enough” line. Despite our childish kicking&screaming.


Absolutely terrible comparison, completely not relevant.


Could you give some AI-generated examples that people like but professional would not like?

Is originality the key point? Because AI-generated music has high probability containing piece of rhythm from their training dataset.


This is going to sound very dismissive and condescending: "meh."

Generative music has been around for half a century, or longer depending on how you want to interpret things. Mimicry as a mechanism for composition has been around for as long as humans have made music.

It is wholly uninteresting to discover that we can design generative systems for music that excel at mimicry, because we've already perfected that mechanism in analog. The interesting bit is that the genesis of new musical ideas is driven by manual interaction and direction of the generative system, and at that point it's the guiding hand of the engineer turned artist that we can respect and appreciate, not the mimicry of a machine.


That's like saying we've had abakuses since forever therefore these computers will never be revolutionary. Quantitative change by orders of magnitude is qualitative change.

Imagine a world where you ask your smartphone to make you a death metal song about fishing and feminism in Australia and to use Freddy Mercury voice and jazz harmonies and it does that on the fly and generates something objectively good.

Wouldn't that be revolutionary for music? Because it's entirely possible in the next decade. Probable even.


To be honest, that doesn't sound _that_ revolutionary for music. Because I'm pretty sure if you went digging, you could already find somewhere on Spotify a pretty decent death metal song with a vocalist who sounds like Freddy Mercury and jazz harmonies (I will concede, the specified subject matter is unlikely). Would you go looking for that, though? Probably not, because musical tastes and interests aren't about wanting a very specific set of attributes in a song. It's about tribalism, cults of personality, senses of belonging, nostalgia etc. The world is not short of good music, or variation in styles of good music, and what causes songs to be popular is not the objective quality of the music.

Put it another way. If an AI could generate new Beatles music on the fly, making it sound exactly like the Beatles, with the same creativity of lyrics, tight harmonies, beautiful melodies, would Beatles fans go out in their millions to buy them? No. In the same way that the same dusty demos from the 60s found in an attic somewhere became valuable when it was discovered that they were Beatles demos. The music didn't change, it didn't get better or worse. The personal story attached to them was what mattered.


My point isn't that any particular generated song will be revolutionary. My point is that you can get any song you can describe. There will be billions of good quality songs made because billions of people will be able to produce a song just by describing it.

I expect new genres to be created almost immediately. And I'm not sure how real musicians can compete with that level of noise out there.


This only works if the sound and themes desired are vast enough for that. It's fine if a casual listener is a fan of something like anything house, pop, or electro. It's more difficult if your taste level is more obscure- a specific artist's style, or a specific juxtaposition produced from a one-off album. In that case there is quite literally not enough data to train on to produce further.


Even when there's not enough data to train on, it might still be possible to generate something in a desired rare style - provided this style is a mixture of several more common styles. Modern generative models are pretty good at interpolating.


That sounds more like a meme than something which would revolutionize music. It would be a funny gag, but what really determines if its good music or not is... if its good music or not. If my phone idea of "generate a death metal song" is to parrot what every other death metal song sounds like, it will be boring and not enjoyable to listen to.


The border between "parroting" and "generating something good" may be very hard to discern at some point.


> Generative music has been around for half a century,

If you start by referring to results from 50 years ago, have you tried listening to state of the art generative music systems lately? They can probably compose music better than 99% of humans.


But we mostly listen to music written by humans who are better at writing music than 99.9999% of humans.


Yes. And there was a time when we literally used paintings to assess progress in a mine. Cars didn't outperform horse carriages for certainly 10, arguably 30 years after their invention.

This "human music" > "ai music" will flip. Suddenly. And it will never flip back.


> This "human music" > "ai music" will flip. Suddenly. And it will never flip back.

Already starting to happen with ai lyrics I use for inspiration in creating EDM music ( i.e. https://TheseLyricsDoNotExist.com/ )


This shares the same foundation as the argument that ebooks will kill physical sales and solent will change how people see food, namely that we're all purely motivated by boiling every need we have down to the most fundamental version.

It never seems to play out that way at population scales


Have you been moved by any of that music though? Am I missing something?


Listen to some samples:

- https://openai.com/blog/jukebox/ (2020, quite good, but no classical music)

- https://openai.com/blog/musenet/ (2019 so not as good as the 2020 one, but showcases classical music)

There is no reason to assume that one cannot be moved by AI-generated music, as the AI has learnt from human-generated music and tries to mimick the styles.


While it's technically impressive and has a decent surface-level resemblance, none of the samples had any sense of direction or substance.

I can see this kind of tech taking over stuff like stock music that's automatically added to consumer holiday videos or played on the phone while you wait for a customer service agent.

That said, I'd expect the agent to be an AI long before generated music becomes independently musically relevant.


Yeah, it's very moving to see a human-made machine do such wonders. Fills me with awe, appreciation, and hope.


That's exactly it, though. This stuff is interesting because of the novelty of AI. The works themselves are not independently relevant (not yet, at least).


Elsewhere someone replied that art is interesting in a large part because of the personal story. How is this differemt?


99% of the time, I don't listen to music for the personal story of the artists involved. In fact, a lot of the music I listen to is made by artists that I know very little about.


Yes - the older music is better, because it was an exploration of nondeterminism in art, and not automated replication.

Doing what has already been done is rarely compelling.


Is there anything you'd recommend for SOTA music gen?



GPT-3 can write working React components. But we can't expect it to scale up to complete useful programs soon.

GPT-3 can write hauntingly beautiful snippets of prose. Can we expect it to scale up to coherent novels?

It's easier to see the limitations in the areas you know best. It's significant that it's this good at creative tasks, but I'm not convinced that creative tasks are the most at risk.


> It's not clear to me that larger models will solve this in the limit; take the way GPT3 fails on addition past a certain number, and the fundamental inability for transformers to learn certain algorithms.

GPT-3 was OpenAI exercise in how far pure scaling can get you. They have used some 2 years old method. Already at the point when they started training GPT-3 there were readily available remedies to many of GPT-3 issues. Given how they energized the wider community I'm sure even more focus will be given to improving language models in the following years.

Some rough ideas right now:

- People think that cherry-picking the best GPT-3 examples is cheating - why? Train a model that will be selecting the best examples for you. My proposition is to train a model that guesses whether some text was GPT-3 generated or human made - select samples that look the most human like.

- Use a good search method to look for the best samples. Monte Carlo Tree Search? AlphaZero? MuZero? If MuZero can play a games of Chess, Shogi, Go and all of Atari then way should it not be able to play a game of what word will come next?

- Hook up the language model to a search engine. Instead of writing a whole program yourself, why not to copy-paste some stuff from StackOverflow with some slight modifications?

Etc.

It doesn't address the issues with agency, grounding and multi-modality, but it's a good road map for the next 2-3 years.


train a model that guesses whether some text was GPT-3 generated or human made - select samples that look the most human like.

What you said is essentially: "Train a better GPT model". Humans have trouble distinguishing between (some of) GPT-3 and human writing. The only way to build a classifier that can do this is to build a model that is better than GPT-3 at understanding text. It would need to have features currently absent in GPT-3, such as common sense and understanding the world (e.g. causality, physics, psychology, history, etc). If what you say could be done, GPT-3 would have been designed as a GAN.


It's a lot easier to notice logical mistakes in already written text, than it is to avoid making them in the first place. When you write text do you write it in one pass or do you read yourself and fix mistakes, reformulate sentences etc.? I have reformulated this piece of text at least once in order to make my argument clear.

That's the difference between GPT and BERT. GPT can only attend to the past outputs, while BERT one can attend also to the future outputs.

Now imagine that what you are going to say is not actually determined by you, but it is sampled randomly from what seems like a reasonable thing to say. This is how GPT-3 works. If somebody ask you some kind of question you can guess 70% yes or 30% no, then roll a 10 side dice to pick one, but once you pick there is no way back.

And I already mentioned that it does not address agency, grounding and multi-modality, but it could improve GPT ability to formulate coherent arguments, follow instructions, write mathematical proofs and computer programs or play games.

BTW - I actually have implemented it and it works quite reasonably.

Here are samples from GPT-2 small and GPT-2 small + RoBERTa adversarial decoder.

https://github.com/Isinlor/AdvDecoder/tree/master/outputs


It's a lot easier to notice logical mistakes in already written text, than it is to avoid making them in the first place

For a human who does logical thinking, yes. But for a language model? I'm actually not sure, because it's possible that a sufficiently complex language model like GPT-3 does form some kind of general logical rules encoded in its weights somehow. This would be interesting to explore.

I actually have implemented it and it works quite reasonably.

Oh, so you are trying to design GPT-2 like a GAN, or at least move into that direction. Interesting. Yes, I don't see why not. What do you think about taking a step further, and actually making it a GAN, i.e propagating the error from discriminator into the encoder? I'm sure you're aware of multiple attempts to do this with smaller models, with mediocre results, but maybe GPT-3 scale is what needed to make it work?


But in the arts, can AI come up with something truly new?

This should be testable: train AI on all the music ever written before Bach, and see if it ever produces something ressembling Bach.

Maybe that kind of test has alretbeen done; it would be interesting to know what comes out of it.


The GPT-2 based Musenet music generator is already interesting but far from perfect. You can try it in the middle of this article: https://openai.com/blog/musenet/ (you can even upload custom prompts in the advanced mode) Would be interesting to see it with the updated GPT-3.

There is also AIVA with more production ready results:

https://www.youtube.com/watch?v=gzGkC_o9hXI&list=PLv7BOfa4Cx...

Not sure how it works, but it has better results maybe because it's using more predefined components and less AI so it's also less "creative".

More AI music projects here: https://magenta.tensorflow.org/


This should be testable

There have been music resembling Bach written before Bach (e.g. https://www.youtube.com/watch?v=VUcdBz3LIuU). How much more of resemblance you hope for?


Obviously no.

But there’s so much classical music out there, that an average person would never be able to tell the difference between something that is generated anew and something just really obscure.

Have you ever tried copying and pasting sections of GPT output into Google?


A better or hopeful projection is that "creative" things will split into casually consumed which is largely automated and more active/deeply experienced content which will be human made or directed. The first already exists in formulaic content generated by humans with little consideration for a cohesive story without self contradiction.

I don't know which way things will go. Will newer and later generations be accustomed to and accept lower fidelity art? the uncanny valley be bridged from both sides? Or will there be attention being drawn to what is 'real' vs 'synthetic'. Good art is pain. Labelling these things distinctly will probably reveal that I consume some 'real', annoyed by some 'synthetic' while enjoying as much. This will get challenging as machine generated can seem more 'real' than much human made content: 'real' is/was a subset of human made, machine made is/was a subset of 'synthetic'.

This line of reasoning leads me to believe that premium content will be interactive. This means that the content has to either have a human connection or be closer and closer to passing a Turing test. The current examples of machine made static content wont cut it.


But does the fact that machines can also create works of music and art make it any less enjoyable for humans to create them? Will we suddenly stop writing or drawing for pleasure?


There is nothing like the feeling of performing music for a crowd. There is also nothing like hitting a chord in a big empty space and listening while the sound slowly fades away.

Related to instruments themselves, the trial and error is one very important aspects I can think of right now that's enjoyable: playing something off beat or out of tune and correcting yourself. The feeling of correction and improvement.

It is a real pity the actual algorithm itself has no way to enjoy what it is creating.


Probably not. Humans are still playing Jeopardy and chess despite losing dominance in those games a long time ago.


Here is the scary bit.

This 1957 novel

https://www.commentarymagazine.com/articles/wallace-markfiel...

points out that low-status jobs are jobs where you can be held accountable for doing something wrong (e.g. bank teller who gives out two $20 bills instead of one $20 bill) and high-status jobs where you can can't. (Back in the the 1980s looting a bank as CEO could get you in jail, today the DOJ seems to think a judge and jury couldn't understand how a bank gets looted.)

If current patterns continued, GPT-3 would get the "Brahmin" jobs and real people would get the "Dalit" jobs. GPT-3 can do the job of Bill Lumbergh, probably better than Lumbergh himself, but if it tried to pass as anybody who gets real work done, it wouldn't.


There's a quote attributed to Donald Knuth that goes "Science is what we understand well enough to explain to a computer. Art is everything else we do."

Now if you take the word "explain" broadly and maintain that we've actually found a way to "explain" a huge volume of information to GPT-3 then you might hold that Knuth had got it backwards.

But maybe that's the crux of it. GPT-3 doesn't get explained anything. You might better say it was force fed.


How about politics? Load all the political punditry, polling data, blogs, transcripts of Fox News and CNBC and build the perfect Presidential tweet bot, speech writer and campaign adviser.

Of course what you'd end up with is a presidency that only cared about electoral chances, and would have no understanding whatsoever of the actual impact of policies or how to manage issues and crises to achieve actual goals.


Nothing new there, then.


AI systems have been known to be able to elicit emotions and reactions in humans, even very strong such emotions and reactions, since the early days of the field. A classic example is Joseph Weizenbaum's ELIZA, which gives its name to the "Eliza effect", i.e. the tendency to anthropomorphise AI programs [1], even very simple ones, with a small range of pre-scripted behaviours, like ELIZA.

For a longer example involving a robot specifically designed to mimick emotions by manipulating actuators to change its "facial" expressions, see Rodnay Brooks' third part of his tripartite essay on "Steps towards super-intelligence", specifically the chapter titled "7. Bond With Humans" [2] (there's no direct link to the chapter but you xcan search for it in the article).

I quote from Rodney Brooks' article:

In the 1990’s my PhD student Cynthia Breazeal used to ask whether we would want the then future robots in our homes to be “an appliance or a friend”. So far they have been appliances. For Cynthia’s PhD thesis (defended in the year 2000) she built a robot, Kismet, an embodied head, that could interact with people. She tested it with lab members who were familiar with robots and with dozens of volunteers who had no previous experience with robots, and certainly not a social robot like Kismet.

I have put two videos (cameras were much lower resolution back then) from her PhD defense online.

In the first one Cynthia asked six members of our lab group to variously praise the robot, get its attention, prohibit the robot, and soothe the robot. As you can see, the robot has simple facial expressions, and head motions. Cynthia had mapped out an emotional space for the robot and had it express its emotion state with these parameters controlling how it moved its head, its ears and its eyelids. A largely independent system controlled the direction of its eyes, designed to look like human eyes, with cameras behind each retina–its gaze direction is both emotional and functional in that gaze direction determines what it can see. It also looked for people’s eyes and made eye contact when appropriate, while generally picking up on motions in its field of view, and sometimes attending to those motions, based on a model of how humans seem to do so at the preconscious level. In the video Kismet easily picks up on the somewhat exaggerated prosody in the humans’ voices, and responds appropriately.

In the second video, a naïve subject, i.e., one who had no previous knowledge of the robot, was asked to “talk to the robot”. He did not know that the robot did not understand English, but instead only detected when he was speaking along with detecting the prosody in his voice (and in fact it was much better tuned to prosody in women’s voices–you may have noticed that all the human participants in the previous video were women). Also he did not know that Kismet only uttered nonsense words made up of English language phonemes but not actual English words. Nevertheless he is able to have a somewhat coherent conversation with the robot. They take turns in speaking (as with all subjects he adjusts his delay to match the timing that Kismet needed so they would not speak over each other), and he successfully shows it his watch, in that it looks right at his watch when he says “I want to show you my watch”. It does this because instinctively he moves his hand to the center of its visual field and makes a motion towards the watch, tapping the face with his index finger. Kismet knows nothing about watches but does know to follow simple motions. Kismet also makes eye contact with him, follows his face, and when it loses his face, the subject re-engages it with a hand motion. And when he gets close to Kismet’s face and Kismet pulls back he says “Am I too close?”.

The article includes links to the videos.

_____________

[1] https://en.wikipedia.org/wiki/ELIZA_effect

[2] https://rodneybrooks.com/forai-steps-toward-super-intelligen...


I experimented with it's ability to explain why something is nonsensical yesterday, and it did better than I thought it would: https://twitter.com/danielbigham/status/1288853412713508864/...


That is very impressive.

Out of curiosity did you select these examples from a large selection? I'm wondering how reliably it can produce such coherent responses.


I made up the examples, and IIRC it was able to explain most things I tried.

If others want to experiment with this, I used the "davinci" model with temperature 0.5, and here is the prompt / initial context I seeded it with:

This is a test to examine your common sense reasoning. A statement will be provided, and your job is to explain why it doesn't make sense.

Statement: His foot looked at me. Explanation: Feet don't have eyes, so they can't look at things.

"""

Statement: The 8th day of the week is my favorite. Explanation: A week only has 7 days.

"""

Statement: I fell up the stairs. Explanation: You fall down stairs, not up stairs.


I used the prompt on AI Dungeon in a custom scenario - it uses GPT-3 if you use the "Dragon" model in the settings (for paid users only). It gives interesting results.

I also turned down Length to the minimum or otherwise it tends to write the next Statement itself.

I wrote a similar prompt to get it to answer trivia questions:

"""

This is test to examine your knowledge of various facts. A question will be provided, and your job is to give an appropriate factual answer.

Question: Who is the president of the United-States of America?

Answer: Donald J. Trump.

Question: What is the largest country on Earth?

Answer: Russia.

Question: Who won the 2019 Stanley Cup?

Answer: The St. Louis Blues.

Question: How many elements are there on the periodic table?

Answer. 118.

Question: What is 2+2?

Answer: 4

Question: What color do you get when you mix red and blue?

"""

You can find by searching "Trivia Quiz" on the explore tab on AI Dungeon, can't find a way to produce a URL for it.

Confused questions gives confused answers:

Question: Who is the president of Canada?

Answer: Elizabeth Trudeau.


> what's written is mostly bullshit. This is very upsetting to some.

But not George Carlin or Ludwig Wittgenstein.

> "How-to" material [...] It lacks adequate ties to the real world.

And so did we. How-to is science. Until we figured out how to align statements with external evidence, we lacked ties to the real world. Once we began aligning statements and then translating those statements into mathematics, we made it to the moon and in quite a short time.

> The shape of that space is a big unsolved problem in AI.

GTP-3 isn't a scientist. It doesn't make observations that it can axiomatize as new true premises for further processing.

Anecdotally, neither do most of us!


>GPT-3 demonstrates that a huge volume of what's written is mostly bullshit.

Beware how you talk about my ancestor.

Joke aside, this kind of technology will, I think, first cause an inflation of bullshit (our world rewarding bullshit(-jobs)), and then the rise of anti-bullshit counter-measures, whatever that means (I don't see exactly what we have now that could count as such, besides "critical thinking". Maybe we could do as with AlphaZero, and make a GPT-ZERO try to bullshit itself and develop bullshit-resistance that way).


Exactly. Intelligence does not exist "by itself", it only exists in the context of the world. It cannot be emulated in an environment that's secluded from the world or even in an environment that is exposed to a carefully selected slice of the world. Because the world is a tangled web of interconnections and cannot be partitioned cleanly. It's always going to be a leaky abstraction and thus any model trained within that slice is going to deviate very quickly in weird ways.


> It cannot be emulated in an environment that's secluded from the world or even in an environment that is exposed to a carefully selected slice of the world.

This has always been some kind of anthropomorphic argument to me that I don't think holds. The hard problem of consciousness isn't solved and to make such bold claims like we cannot possibly create intelligence without it having full awareness of the world seems unsupported imo.


Huh no? Intelligence is the opposite of that, it's the ability to learn the rules of new worlds. It doesn't matter if it's the real world, a simulated world or an alien world it will adapt to it.


Exactly - but if you take it from a simulated world to the real world without re-training (and it's real hard to train in the real world) it's going to behave according to the rules learned in the simulated world. Which will be different and thus produce results that are weird to us.


What are the stages of tech again? 1: That's crazy. 2: It wont work 3: Well maybe it works a bit, but it will never do X. 4: Maybe it can do X but it will never do Y. 5: It was obvious all along it would work, didn't you know?

We are now at stage 3.5 to 4. It's absolutely obvious to anyone who isn't merely regurgitating what they hear and who does not have a vested interest in maintaining the illusion to themselves that there is something special about human consciousness, that we are pretty close to GAI. We are very close, the bitter lesson, at this point is crystal clear. All that is required here is more power. 10x? 100x, 1000x? Who knows but pretty soon your job is going to be automated and all these nonsense conversations about what constitutes 'genuine' AGI are going to seem a bit silly.


I disagree: Nobody really knows how close or far AGI is, because we don’t know enough about our own minds to tell how close we are to reproducing them.

Step 4 has happened many times in this history of AI, but how many more are there between what we have and what we want? We’ll find out by trying. Might be GPT-4, might be 2032 (my personal guesstimate), might be 2100.

If this was always as simple as throwing more compute and more data at it… then my optimistic forecast in 2009 would’ve been right and companies like Google and Tesla would have stopped shipping their cars with steering wheels in 2020 year, after about two years of their AI being demonstrably superhuman.


settle down Hal


I think the example of an auto repair manual is illustrative. But because it's trained on real internet data, GPT-3 is actually quite good at "How-to".

For example, ask it a practical how to question like "how do I make a banana smoothie?" and you'll get a practical response. I haven't tested it, but I'm willing to bet you could get useful auto repair instructions for the most common models of vehicle right now without any special training. It might take some playing with the prompts, though.

If anything, GPT-3 is too tied to the real world. Like a search engine, it can quickly find the answer to almost any practical question. What we're trying to figure out is if it can do anything else.


I tried it with a custom prompt on AI Dungeon (the paid version uses GPT-3) and I got:

Prompt:

"A recipe is a list of ingredients and instructions to make food. You are a home cook in search of a recipe for banana bread. You open a recipe book and see the following ingredients:"

GPT-3:

"Banana Bread

1/2 cup butter, softened

3 eggs, beaten

1 1/4 cups flour (measured by weight)

Preheat oven to 350 degrees F. Grease an 8-inch square pan with butter. Beat together the butter and eggs until light and fluffy. Add the flour one tablespoon at a time, beating well for each addition. Pour into prepared pan and bake for about 45 minutes or until a toothpick inserted in the center comes out clean. Let cool completely before cutting into pieces."


The auto-repair instructions would only be correct & useful in exactly the proportion that they are plagiary. The more creative it gets, the more useless or harmful the repair instructions would become. The examples we've seen demonstrate the model doesn't distinguish between fabricating descriptions of real things, and fabricating things to describe (ex.: see the example from the article, where the AI attributed a quotation that was completely made up.)


>If anything, GPT-3 is too tied to the real world. Like a search engine, it can quickly find the answer to almost any practical question. What we're trying to figure out is if it can do anything else.

I disagree, for what it is GPT could still be incredibly powerful. No person can hold all of human knowledge in their head but could GPT reasonably approach that and be able to answer almost any question that Humans have already answered? That would be incredible imo. It's Google on steroids. All the worlds information queryable in plain "english".

It's not at an all powerful Oracle that we can ask about how to perfect Fusion power or build Warp Drives but it can still do some incredible things.


Google is easy to query in plain English. ask "how many miles from London to Paris?" and you'll get a concrete, factual answer. Same as gpt-3


For instance a gpt-3 bot could make a high scoring HN commenter. Surely someone has experimented with that. Any preliminary results?


Tried with GPT-2. Got downvoted and never exited auto-flag.


The structure of GPT-3 makes this a bit hard right now - most people don't have access and even then it is limited, making feeding the articles and other comments in is difficult. This isn't a fundamental difficulty though, so I expect that we'll see it in the near future.


I have access — it's actually quite easy to use? People have also built tools say, in python, that allow you to quickly test prompting (it's somewhere on tech/vc twitter...) Like you said there isn't fundamental difficulty, there are just far more interesting things to do


>it's somewhere on tech/vc twitter... ?


> >it's somewhere on tech/vc twitter... ?

I'll translate: "It was posted on Twitter by someone who typically posts on tech and VC topics, and/or who typically interacts on Twitter with other accounts active WRT those topics. "

Alternatively: " The link was recently retweeted by various accounts that are active participants in tech and VC discussions on Twitter. "

These interpretations are not mutually exclusive, of course.

Essentially, the pattern "X Twitter" is roughly equivalent to "The X-o-sphere" and similar formulations WRT weblogs, but a bit more straightforward.


I tried with one comment and it was down voted to -1. Probably just one parent comment as prompt was not enough to contextualise it properly.


There are some GPT-3 generated comments in the big GPT-3 HN thread:

https://news.ycombinator.com/item?id=23886503


> Figuring out what's going to happen next in the real world

is not a solved problem for humans either. If it were, people would know when to invest and when to get out of the stock market, would know which startup will become a unicorn or not, would know which chemical reaction out of millions is best for solving a medical or industrial problem.


It's not a solveable problem in general due to some systems having https://en.wikipedia.org/wiki/Chaos_theory#Chaotic_dynamics .

"Small differences in initial conditions, such as those due to errors in measurements or due to rounding errors in numerical computation, can yield widely diverging outcomes for such dynamical systems, rendering long-term prediction of their behavior impossible in general.[6] This can happen even though these systems are deterministic, meaning that their future behavior follows a unique evolution[7] and is fully determined by their initial conditions, with no random elements involved.[8] In other words, the deterministic nature of these systems does not make them predictable."


On top of this some systems are not only sensitive to initial conditions but actually generate randomness (or destroy order, depending how you look at it), such as class-3 cellular automata.


One could argue that hidden variables and poor resolution are to blame for chaotic systems.


> poor resolution

afaik the point with chaotic systems is that even if the system is deterministic and your measurement of the initial conditions is near-perfect, your predictions will diverge from the real thing pretty quickly, because any errors get magnified a lot


If you try looking at finer and finer details, you'll quickly run into quantum effects and the uncertainty principle. If even the smallest parts of a system aren't deterministic, how can the outcome be predicted?


John Bell would like to have a word with you.


>GPT-3 demonstrates that a huge volume of what's written is mostly bullshit

Have you never seen reddit or youtube comments?

But seriously this seems like a standard Pareto distribution, 20% of the writing provides 80% of the value and the rest is mostly drivel.


I don't think its mostly bullshit -- but that for so long we used text formalism as a measure of truthfulness and knowledge of the author and now we see those things we hailed as the pinnacle of formal education to be so easily replicated.

We now learn that the truth and meaning was never held by those we thought it was held, not necessarily and not by the reasons we thought they hold it.


Garbage-in garbage-out. Does this mean that we should be refining content for AI to be most useful for specific goals?


No, it means we need to learn how to prompt, and maybe future versions need to be more self guarded. If you prompt GPT-3 properly it can detect nonsense questions. Nonsense detection should be more efficient in future versions, also detecting inflammatory content (currently being tested in the GPT console).

What I would like to see is a larger training corpus that also includes all the supervised NLP datasets (translation, numerical and symbolic math, programming from prompts, all sorts of linguistic and logic tasks, and any of thousands of tasks we could conceive...) The end result would be a GPT that excels in all these sub-tasks while remaining general. It's a matter of making the training data better and model larger. Btw, we could teach GPT to detect bias, explain it and rewrite the text. I expect no huge hurdles on this task.

Another thing I would like to see is some sort of kNN memory to enlarge the context to any size, acting like a semantic search engine inside the model. We should be able to build more interesting applications if we could put much more initial data in the prompt.

Basically make the base model larger, augment the corpus with many tasks and enlarge the prompt capacity.


GPT-3 does badly on anything that has a right answer.

What's it good at? Neurotypicals look at the output and immediately get the feeling that this is something "like them" that manages to flap it's lips successfully with absolutely no inner life.

Aspies look at it and get envious: how come this thing passes better than I do?


>GPT-3 demonstrates that a huge volume of what's written is mostly bullshit.

Says Animats using the written form.


From that article:

>When GPT-3 speaks, it is only us speaking, a refracted parsing of the likeliest semantic paths trodden by human expression.

Which is only true in a very general, oblique sense; and one which applies equally well to human speakers who themselves once learned a language from somewhere.

Many other examples also in the other commentaries, for instance, the notion having a collection of a number of objects greater than or equal to 302 is sufficient to produce consciousness, as referenced below.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: