We Want to Text 300 People's Partners an AI Love Poem. What Do You Think Happens?
Here is an experiment we want to run, described before we run it, so nobody can accuse us of writing the conclusion first.
Two groups text their partner something affectionate, unprompted. One sends a genuinely good love poem written by a top-tier chatbot: well-metered, specific, moving. The other sends a five-second doodle of a lopsided heart with a typo in the caption.
Then we look at what comes back.
We have a guess. We want yours first.
What gets the warmer reply?Thanks, your pick is counted.
Why We Think This Is Worth Testing
The interesting question is not whether AI can write a good poem. It plainly can, better-metered and more coherent than most of us manage unassisted, on a first pass, at 2am.
The question is whether quality is what a message like that measures. An unprompted romantic text is not really an artifact. It is evidence. It says someone thought about you and spent something to say so. If the spending falls to zero, it is not clear the evidence survives, however gorgeous the object.
Our hypothesis: recipients will respond more warmly to the worse artifact, because the badness is the proof. A wobbly heart with a typo could not have come from anywhere else.
We could easily be wrong. Plausible alternatives:
- Nobody notices. People receive a nice message, feel nice, and never wonder where it came from. Quite likely, honestly.
- The poem wins. Effort gets inferred from length and polish, not crudeness, and a long beautiful poem reads as more effortful than a scribble.
- It depends entirely on the couple. Relationship length, age, and how the two already talk swamp the effect.
If the poem wins, we will publish that. A result that contradicts our own marketing makes a better post than one that flatters it.
How We'd Run It
Three groups, not two. The two-arm version confounds several variables at once: the doodle differs from the poem in medium, length, effort, and surprise simultaneously, so a result would not tell you which mattered. A third arm fixes it:
- AI poem generated, unedited, sent as-is.
- Human short text the sender writes one affectionate line themselves.
- Doodle five seconds, drawn by hand, no redraws.
If the doodle beats the human text, the medium matters. If they tie and both beat the poem, what matters is that a person made it, not that it was drawn. Either way, a more interesting finding than we started with.
Around 100 per group, not 500. Recruiting a thousand people is harder than it sounds, and 300 is ample to see a large effect. Better a completed study of 300 than an abandoned one of 1,000 dying in a spreadsheet.
Pre-defined measures, fixed now rather than after we have read the replies and grown attached to a story:
- Time to reply
- Reply length
- Presence of affection markers (terms of endearment, emoji, kisses)
- Whether the recipient questions the message's origin
- A one-question follow-up to the partner, 24 hours later: how did that message make you feel, 1-5
That last one matters most, and it is the measure the original concept lacked entirely: the recipient's actual experience, rather than our reading of their text.
The Part We'd Have to Be Careful About
This involves sending something slightly deceptive to someone who did not sign up for it, and that someone is a person the participant is in a relationship with.
For most couples, harmless fun. For some, not. "You sent me an AI-written love poem as part of a marketing experiment" could land like a brick in a relationship already under strain, and no blog post is worth doing that to somebody.
So the rules:
- Participants debrief their partner within an hour. No exceptions; the debrief is part of the task, not an afterthought.
- Anyone whose relationship is in a difficult patch is asked not to take part. Stated plainly at signup.
- The 24-hour partner survey is sent after the debrief, so the partner knows what they are rating and consents to their answer being counted.
- No message content published without both people agreeing.
Worth saying aloud rather than quietly handling, because an experiment about sincerity that treats people's relationships as raw material would be its own kind of joke, and not a good one.
Known Weaknesses
Stated now rather than pointed out later by someone smug:
- Our participants will self-select, and people who volunteer for an experiment run by a drawing platform are not a neutral sample. We will say so and will not generalise to the population.
- Being in a study changes behaviour. Senders know they are participating, which shapes how they send.
- Reply warmth is a rough proxy for anything meaningful about a relationship.
- One message is a snapshot. It says nothing about what a year of either habit would do.
A fun experiment, not a paper. We will present it as one.
Call It Now
Vote above. We will publish the results, whatever they are, alongside the predictions from this post.
If you want to be one of the 300, sign up here.
What gets the warmer reply?Thanks, your pick is counted.