DESIGN EXPERIMENT

Evaluating AI Synthesis Tools

IDEO U's AI x Design Thinking course walks through this exercise using ChatGPT. After finishing the assigned prompts, I wanted to know if the answers held up — so I reran the same prompts through Claude and Notebook too, to see where they agreed, diverged, and where a person still had to make the call.

The design question

How might we help kids make better decisions around food?

Why compare tools?

Every AI tool sounds confident, but should be used responsibly.

What's below

My working notes, tool comparison, and learnings from the experiment.

Should I use open or grounded synthesis?

Intention and data sources matter. If you are looking to do research on specific data and are not risk-tolerant, use grounded. If you want to be more creative and are willing to do more audits, use open (with caution). When I tested in this experiment, I used the default for each tool.

open synthesis

Claude and ChatGPT typically reason freely over whatever's pasted into the conversation. It's fast and flexible, but if you don't specifically set boundaries and data sources, it can introduce outside assumptions.

grounded synthesis

Notebook only works from sources you upload, and traces each claim back to them. Grounded synthesis is slower to set up, but minimizes bias and lowers risk of hallucinations.

Research source material for all three tools

Synthetic sample data from IDEO, with a quantitative survey of 25 kids (eating habits, household income, produce access) and 5 interview transcripts on food preferences and family mealtime.

Notes & observations

The exercise in this class had 6 prompts for Quantitative and 4 prompts for Qualitative. I took notes throughout the process to compare each tool; below are observations from the first step in each phase. 

This is a meta-synthesis: a synthesis of 3 synthesized research projects in each tool.

Quantitative notes

Quantitative synthesis — Each tool's response from the first prompt was eye-opening, and demonstrated how each tool presented findings

Qualitative notes

Qualitative synthesis — At this stage, Notebook was the most useful, providing a more thorough view of the themes that made it easier to parse and decide where to probe. But there were a few surprises!

Synthesis notes

Cross-cutting synthesis — quick view of what worked and didn't work

What I learned

Each tool offers grounded synthesis, if specified.

I used the default that each tool offered, but am interested in experimenting with the grounded synthesis capabilities of Claude. I was disappointed in ChatGPT, and am not interested in testing it further.

As expected, grounded synthesis reduced risk.

Grounded synthesis made me feel more secure in the results presented, but still needs human review.

Each tool was more different than I expected.

I didn't expect to see such a divergence in results across tools; accuracy, style, and recommended next steps.

You might need more than one AI tool for the job.

Experiment with running two tools in parallel. Convergence is cross-validation.