Recreating Minecraft Is Not a Benchmark

(kuber.studio)

13 points | by kuberwastaken 2 hours ago

4 comments

  • hombre_fatal 25 minutes ago
    > That’s the problem, these tests can’t tell you how good a model is anymore because it’s trivial for labs to optimise for exactly these tests by the next release.

    But they don't prove the claim. Are the models amazing at recreating Minecraft, but the second you swap the word Minecraft out with another game or a custom game, it shits the bed?

    That's not what I see. My feed is full of people using Astra to recreate all sorts of games from Diablo to some random idea they came up with, in ridiculously polished detail like animations that would have taken me weeks of iteration in gpt-5.6-sol but it was a single shot by Astra.

    • enraged_camel 1 minute ago
      You are falling for selection bias.

      For every person using Astra to create a game from some random idea and sharing the impressive result, there are an unknown number who have tried the same thing and gave up in frustration.

    • tripleee 9 minutes ago
      > ridiculously polished detail

      Can you point me to one of these? The only Astra game I tried was the incredibly janky Mario Kart clone from OpenAI where you could fly by spamming jump

  • johnsonjo 28 minutes ago
    Though I somewhat think some benchmarks are silly like the article says. I saw someone on YouTube recently take a picture of a building across the street from them (seemed like it was in NYC), and asked GPT 6 Astra to make it in blender. It did a surprisingly good job in 30 minutes. So though these benchmarks don't seem to mean much you could always add a touch of randomness to them like the person in the YouTube video did, but the problem with that is how would you compare the benchmarks in any clear way if they aren't even consistent? Regardless it appears LLMs are getting this good at the general task and not just at the particular instances of said task.
  • Kuinox 33 minutes ago
    I cant select text nor click links on this page with firefox.
    • chuckadams 24 minutes ago
      Running Firefox here, no problems whatsoever even with UBO and Privacy Badger disabled (the twitter embeds get blocked at the DNS level, but I doubt those are the problem).
      • Kuinox 21 minutes ago
        Works on phone but not on my linux desktop.
        • chuckadams 18 minutes ago
          Mac here. There's nothing all that interesting going on with the JS on that page, so I would suspect a bug with Firefox and/or your desktop environment.
  • alephnerd 26 minutes ago
    Most of my and my peers PortCos run their own eval and benchmark sets, simply because they know what they need best.

    The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now.

    This has been the operating assumption for me and my peers, and has largely played out that way.

    That said, this has always been an issue with benchmarking since the very beginning. DB Benchmarks, compute benchmarks, and others that were external facing were always inherently a content and product marketing tool. The actual internal benchmarking used to model, understand, and enhance your product was always a closely held secret.

    Most of these conversations are happening, but largely in person and not on HN.