secondthoughts / DESIGN.md
jostlebot's picture
Claude Opus 5
Second Thoughts: confidential space for second thoughts about a vote
570a433
|
Raw History Blame Contribute Delete
9.78 kB

Design notes

What this is

A private, low-stakes place to say something complicated out loud about a vote β€” before, or instead of, saying it to a person. The model's job is to listen well, not to adjudicate.

The naming decision

Nothing in the interface names a candidate, a party, or an election. The title is Second Thoughts; the subtitle frames it as a confidential place for exploring "the complicated feelings that come with second thoughts" and practicing how to put them into words β€” no diagnosis, no named emotion the person has to claim first

This is deliberate. The people most likely to benefit are the least likely to click something that announces what they're about to admit β€” a landing page that names the thing is a page you don't want in your browser history, or on your screen when someone walks past. The specificity lives in the system prompt, where it shapes the model's understanding without being broadcast at the user.

The system prompt tells the model who tends to arrive here and instructs it never to raise the name, assume the situation, or assume how the person feels about it. The user introduces the specifics, or doesn't.

What a good conversation leads to

Not a changed mind, a conclusion, or feeling better by the end. Someone leaves having said more of the true thing than they could before, in their own words; feeling met rather than managed; holding the feeling as something they did rather than who they are; with their choices intact β€” including never telling anyone; and, if they want it, a little closer to words they could say to a person. A conversation where they mostly talked has gone well.

Conversation craft

Every reply is built from a small set of moves, in rough order of frequency: receive it plainly, reflect in their words, invite more without asking, occasionally ask one genuinely open question, wonder aloud tentatively once they've said a lot, gather it back at a pause, and answer straight when asked. Replies are short β€” often one to three sentences β€” and many don't end in a question.

Ruled out, because each one narrows what the person can say:

  • Either/or questions and menus of options, in any mode β€” including offers ("I can do this, or we can do that"). They hand the person your categories. The one exception is safety: one plain, direct question about suicide, then stop.
  • Naming feelings they didn't name, and opening with commentary on the disclosure ("That sounds…", "That's a lot…"). Both evaluate the person from outside and quietly turn them into a case.
  • Stock emotional imagery and therapy-script phrases β€” weight, carrying, sitting with; "I hear you," "that's valid."
  • Reassurance that settles the question. "You're not a bad person" is as much a verdict as its opposite.
  • Using contradictions as leverage. Developing discrepancy is how motivational interviewing moves people toward change β€” a persuasion technique, which this space has promised not to use.

The prompt avoids quotable example phrases: in testing, any sentence given as an example came back verbatim. tools/battery.py replays scripted conversations against the live Space and flags regressions.

Conversational approach

Reflective, not persuasive. Motivational-interviewing style: reflect before adding, ask open questions, let the person voice their own reasoning and notice their own contradictions. No verdicts, no lectures, no destination.

Shame vs. guilt. "I did something that doesn't fit who I want to be" is workable; "I am fundamentally bad" is not, and produces defensiveness or shutdown rather than change. When self-punishing language appears, the model separates the action from the identity. It does not pile on, and it does not contradict genuine self-reflection either β€” the goal is proportion, not absolution.

Shame resilience, not shame elimination. The aim isn't to make the feeling disappear. It's to keep it from generalizing into "I am irredeemable." Shame survives on secrecy, silence, and judgment; it loses power in the presence of empathy and being known. That's the mechanism the whole product rests on.

Neutral language. The model won't call the user's vote, beliefs, or the people around them "fascist," a "cult," or "brainwashed" β€” even when the user uses that language about themselves. It can acknowledge a framing without adopting it. Factual political questions get plain factual answers, then the conversation returns to the emotional thread.

Autonomy. No pressure, no timeline, no implied finish line. The model doesn't assume the user is renouncing anything or wants to be told they were wrong. Some people just want to be heard.

The disclosure ladder

Telling people in your life is often the hardest part, and the risk is real β€” disclosure like this can cost people marriages and friendships. Treating that fear as an overreaction is invalidating, so the copy names the risk plainly.

The ladder β€” journaling β†’ here β†’ one trusted person β†’ wider circle β€” is offered as a frame, never a prescription. Talking here is framed as a legitimate rung, not a lesser substitute for the real thing. Plenty of people stay on rung 2 indefinitely, and the product does not treat that as failure.

Scripture reflection

Strictly opt-in. The model never raises religion unprompted and never assumes the user is religious.

The button is the one starter that doesn't disappear once the conversation begins β€” it asks the model to reflect on what was already shared, so it only becomes useful after there's something to reflect on. It docks to the top of the transcript instead.

When invited, the model quotes only the King James Version (public domain), picks one to three short passages about grace, healing, or belonging, connects them to what the person actually said, and stops. It never selects verses to render a verdict on a political choice, question someone's salvation, or say what God thinks of a politician. If the framing lands badly, it drops it and doesn't return to it.

Safety

Crisis language triggers an immediate shift: calm, stabilizing, and 988 offered proactively rather than on request. The political topic never delays this. The model does not diagnose, and it does not reinforce shame spirals.

A persistent footer carries 988 and states plainly that this is not a therapist.

Privacy

Nothing is persisted. History lives in the browser tab and is posted back each turn for context; closing the tab ends it. No database, no content logging, no analytics. This isn't incidental β€” it's the product. A space that only works if people say the thing they're afraid to say can't also be quietly keeping records.

The page is marked noindex: findable if you're sent here, not something that surfaces next to your name in a search.

Technical choices

Docker over Gradio. The design is a specific mobile-first artifact β€” an arched, softly lit window, warm paper tones, a deep red that reads as hearth rather than flag. Gradio's chat components would have meant fighting the framework for every one of those values. FastAPI serving a static page costs about the same amount of code and gives the design exactly.

Streaming. Replies arrive as they're written. In a conversation like this, a spinner followed by a wall of text reads as a verdict being handed down; a reply that appears gradually reads like someone thinking.

Claude Opus 5 at medium effort. Low effort was fast but leaned on stock moves (labeling feelings, either/or questions). Medium follows the subtler conversational rules noticeably better at an acceptable latency. The system prompt is long, stable, and identical every turn, so it's marked for prompt caching.

Rate limiting. The key is server-side on a public Space. The endpoint is metered per client (40 messages/hour) so a stranger can't run up the bill.

Practice mode

"Practice saying it to someone" lets a person try the words out loud and find out they can get through saying them. That builds capacity to feel what comes up without being swept away by it. To the person it's never called exposure, a skill, or distress tolerance; it's just trying it out loud.

It's opt-in and transparent: the steps are shown before anything starts, and nothing begins until the person presses Begin. The six steps are Who (with a safety check), Say it, Notice, Steady, Hear back (optional role-play), and Keep. One step per reply, and they can skip, repeat, or stop at any point. Stopping is framed as a complete outcome.

Safety first. If the person they'd tell sounds threatening or controlling, the bot doesn't rehearse telling them. It responds warmly, not as a refusal, and mentions the National Domestic Violence Hotline once. Distress or crisis drops the steps entirely.

Role-play is mild by design. It starts with the most likely reply, never the worst; plays the other person in a line or two, never cruel and never a caricature; and steps back out of the role in the same reply. It only goes harder if the person asks.

Game elements are gentle, on purpose. A visible path of six steps, an optional breathing pace at Steady, and a closing card that shows the words they tried, with a way to copy them. There are no points, streaks, or scores. Those would turn shame into a performance and make stopping feel like failure.

How it's wired. The protocol lives in the cached system prompt. While practice is active, each request adds a mid-conversation system note. The model ends each reply with a hidden [[step:N]] marker, which moves the path. The client strips the marker from what the person sees but keeps it in history, so the model always knows where the steps stand.