I ran a similar test last month, splitting research across Claude and Perplexity to see which one caught more contradictions between sources. Claude was better at synthesizing a coherent POV, but Perplexity surfaced more disagreeing sources in the first place, which matters more for the cross-checking step you described. Curious if you tried a second model as a sanity check on the first one's output.
I ran a similar test last month, splitting research across Claude and Perplexity to see which one caught more contradictions between sources. Claude was better at synthesizing a coherent POV, but Perplexity surfaced more disagreeing sources in the first place, which matters more for the cross-checking step you described. Curious if you tried a second model as a sanity check on the first one's output.
No, we didn't for this one. Used different ChatGPT models for this experiment. But thanks for telling us; this sounds interesting.