IDEO U's AI x Design Thinking course walks through this exercise using ChatGPT. After finishing the assigned prompts, I wanted to know if the answers held up — so I reran the same prompts through Claude and Notebook too, to see where they agreed, diverged, and where a person still had to make the call.
How might we help kids make better decisions around food?
Every AI tool sounds confident, but should be used responsibly.
My working notes, tool comparison, and learnings from the experiment.
Intention and data sources matter. If you are looking to do research on specific data and are not risk-tolerant, use grounded. If you want to be more creative and are willing to do more audits, use open (with caution). When I tested in this experiment, I used the default for each tool.
Claude and ChatGPT typically reason freely over whatever's pasted into the conversation. It's fast and flexible, but if you don't specifically set boundaries and data sources, it can introduce outside assumptions.
Notebook only works from sources you upload, and traces each claim back to them. Grounded synthesis is slower to set up, but minimizes bias and lowers risk of hallucinations.
Synthetic sample data from IDEO, with a quantitative survey of 25 kids (eating habits, household income, produce access) and 5 interview transcripts on food preferences and family mealtime.
The exercise in this class had 6 prompts for Quantitative and 4 prompts for Qualitative. I took notes throughout the process to compare each tool; below are observations from the first step in each phase.
This is a meta-synthesis: a synthesis of 3 synthesized research projects in each tool.

Quantitative synthesis — Each tool's response from the first prompt was eye-opening, and demonstrated how each tool presented findings

Qualitative synthesis — At this stage, Notebook was the most useful, providing a more thorough view of the themes that made it easier to parse and decide where to probe. But there were a few surprises!

Cross-cutting synthesis — quick view of what worked and didn't work
Each tool offers grounded synthesis, if specified.
I used the default that each tool offered, but am interested in experimenting with the grounded synthesis capabilities of Claude. I was disappointed in ChatGPT, and am not interested in testing it further.
As expected, grounded synthesis reduced risk.
Grounded synthesis made me feel more secure in the results presented, but still needs human review.
Each tool was more different than I expected.
I didn't expect to see such a divergence in results across tools; accuracy, style, and recommended next steps.
You might need more than one AI tool for the job.
Experiment with running two tools in parallel. Convergence is cross-validation.