Hacker News (curated)new | past | comments | ask | show | jobs| show hidden

Edit: I should have read through the whole thing first, ignore me


From the report:

> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.

> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.


I appreciate the response, I should have finished reading through the whole thing first. My initial reaction assumed far less usage of AI to analyze the data.


TFA says as much, and METR said so themselves



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact | github