Using Opus 5.5 to discover a new eyewitness record of the dodo

(resobscura.substack.com)

58 points | by benbreen 4 hours ago

4 comments

  • nl 1 hour ago
    > Epistemological weirdness

    > They are also notably bad at judging the historical significance of what they find.

    I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.

    It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.

    > seven chord groups

    This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.

    • Waterluvian 19 minutes ago
      This analogy may be too close to the real thing to work, but it reminds me of a Chinese room type situation where its entire understanding of the world is through messages of text.

      You say that’s an error a human couldn’t do, but imagine if the human has never seen or touched the kind of item you were making and relied entirely on text descriptions to build its ontology. Off by 90 seems like such a believable mistake.

    • NewJazz 37 minutes ago
      It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.

      Maybe the model isn't intelligence in any form, except perhaps as an imperfect reflection of the intelligence of its training data.

    • morpheos137 16 minutes ago
      In general llms are weak with spatial reasoning. This seems to be an unsolved problem. Probably because human language is generally imprecise spatially and humans think about spatial problems in visual terms. I wonder if having an llm make a 3d design in a format an image model could check would result in a better outcome?
  • early_exit 7 minutes ago
    Good read! How many pages of text were in scope? I'm not sure if the 1615+1629 pages were the total or just a subagent.

    If they were the total I would say it was arguably more impressive the author was able to narrow it down to just 3000 pages than it was to find the dodo mention amongst those!

  • dgellow 2 hours ago
    What a great read, I generally associate substack with verbose, low quality content, but definitely not the case here!
  • deepinquiry 1 hour ago
    [flagged]