Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

but have Claude actually fact-checked it or just provided an opinion?


Actually it was Fabel, and it totally tore apart Gemini's response.


What if you use OpenAI as a judge and check both responses? Not defending Gemini, just curious if it was that wrong, or Fable - that aggressive.


"Lazy" evaluation




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: