I know you're talking about Batty's tears in the rain speech, but I also found the Voight Kampff test pretty memorable:
"The tortoise lays on his back, his belly baking in the hot sun, beating its legs trying to turn itself over, but it can’t, Leon, not without your help. But you’re not helping…. Why is that, Leon?"
> “Claude first generated about 650 unsuccessful ideas. After being told to keep trying, it spent roughly a day and a half orchestrating around 60 Claude subagents, producing 31 million output tokens, running about 2,400 shell commands, writing hundreds of Python scripts, performing thousands of numerical checks and having agents critique each other's work. Human intervention was mostly encouragement rather than mathematical guidance.
In other words, this is a cosmically large amount of trial and error trying basically everything the available data told Claude would be worth trying. Humans would never do this, we are too limited resource wise, so this is a big result in itself.”
Isn’t this similar to how the human brain operates, though? Not the conscious experience you’re having as the creative person, but all the work happening under the hood when your brain chews through a hard problem in the back of your mind until it finally delivers a compelling solution to you in the shower.
What feels like de novo creativity to our personal consciousness might just be the experience of becoming aware of the “final product” once all the postulating, theorizing, testing, and checking has been performed outside of (and prior to) our awareness.
On a different note. I gave a toast at a wedding last month, and at one point in my preparations I asked ChatGPT to brainstorm one-liners. Many were garbage, but one of them ended up in the speech and got the biggest laugh of the whole night. Interestingly, it didn’t come out with the joke out of nowhere, but it took two of my existing jokes and combined them in a novel and pithy way that was way, way better with essentially perfect comedic timing and pace. I doubt the same process could have generated the George Foreman joke, but I still found it creative in its own sort of way.
yes certainly those types of creativity are present in people and models, and can produce a particular type of good stuff
on the joke, did chatgpt realise that was the best one in some sense, or it did it need your experience to tell that was better than the others? or did you not even know yourself it was the best one until you got the laughs?
Great question- no it just served it up among dozens of other options. But it was so clearly head and shoulders ahead of the others (per my perception when I read it; the model didn't seem to be aware that it was the best) that I immediately put it into the speech (and it killed, as predicted). So you can definitely classify it as an outcome of good luck and statistics, but nevertheless it amazed me how many standard deviations above average it could get to.
Another thing I noticed is that compared to a year ago (roughly the last time I asked ChatGPT to help me brainstorm jokes), the floor level of the very worst jokes has gotten much higher. A year ago the worst jokes were gobbledegook that didn't even make grammatical sense, whereas now the worst jokes are mediocre puns that at least clock as jokes even if they're not especially funny. Not difficult to imagine LLMs being consistently funny in 2-3 years.
I know you're talking about Batty's tears in the rain speech, but I also found the Voight Kampff test pretty memorable:
"The tortoise lays on his back, his belly baking in the hot sun, beating its legs trying to turn itself over, but it can’t, Leon, not without your help. But you’re not helping…. Why is that, Leon?"
I actually do need to watch it again
What is it like to hold the hand of someone you love? Interlinked.
> “Claude first generated about 650 unsuccessful ideas. After being told to keep trying, it spent roughly a day and a half orchestrating around 60 Claude subagents, producing 31 million output tokens, running about 2,400 shell commands, writing hundreds of Python scripts, performing thousands of numerical checks and having agents critique each other's work. Human intervention was mostly encouragement rather than mathematical guidance.
In other words, this is a cosmically large amount of trial and error trying basically everything the available data told Claude would be worth trying. Humans would never do this, we are too limited resource wise, so this is a big result in itself.”
Isn’t this similar to how the human brain operates, though? Not the conscious experience you’re having as the creative person, but all the work happening under the hood when your brain chews through a hard problem in the back of your mind until it finally delivers a compelling solution to you in the shower.
What feels like de novo creativity to our personal consciousness might just be the experience of becoming aware of the “final product” once all the postulating, theorizing, testing, and checking has been performed outside of (and prior to) our awareness.
On a different note. I gave a toast at a wedding last month, and at one point in my preparations I asked ChatGPT to brainstorm one-liners. Many were garbage, but one of them ended up in the speech and got the biggest laugh of the whole night. Interestingly, it didn’t come out with the joke out of nowhere, but it took two of my existing jokes and combined them in a novel and pithy way that was way, way better with essentially perfect comedic timing and pace. I doubt the same process could have generated the George Foreman joke, but I still found it creative in its own sort of way.
yes certainly those types of creativity are present in people and models, and can produce a particular type of good stuff
on the joke, did chatgpt realise that was the best one in some sense, or it did it need your experience to tell that was better than the others? or did you not even know yourself it was the best one until you got the laughs?
Great question- no it just served it up among dozens of other options. But it was so clearly head and shoulders ahead of the others (per my perception when I read it; the model didn't seem to be aware that it was the best) that I immediately put it into the speech (and it killed, as predicted). So you can definitely classify it as an outcome of good luck and statistics, but nevertheless it amazed me how many standard deviations above average it could get to.
Another thing I noticed is that compared to a year ago (roughly the last time I asked ChatGPT to help me brainstorm jokes), the floor level of the very worst jokes has gotten much higher. A year ago the worst jokes were gobbledegook that didn't even make grammatical sense, whereas now the worst jokes are mediocre puns that at least clock as jokes even if they're not especially funny. Not difficult to imagine LLMs being consistently funny in 2-3 years.
Fantastic read Stefan!