OpenAI's GPT-6 Astra on ARC-AGI-3

(arcprize.org)

135 points | by vignesh_warar 3 hours ago

12 comments

  • malfist 2 hours ago
    Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
    • matherial 1 hour ago
      It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles.

      I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.

      Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.

      • pavlov 1 hour ago
        When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there.

        It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.

        • phainopepla2 31 minutes ago
          At the risk of sounding like one of those people at Mensa that annoyed you...

          The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.

        • guelo 58 minutes ago
          What does having value or interest to you have to do with intelligence?
          • itishappy 28 minutes ago
            What does IQ have to do with intelligence?

            (Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)

        • sfblah 37 minutes ago
          To be fair, Mensa is a Venn diagram between IQ and being a douche.
      • tintor 1 hour ago
        Pure software benchmarks might be getting saturated, but physical ones aren't.

        Let LLM control a physical robot to perform tasks that average human can do.

      • skybrian 58 minutes ago
        Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.
      • ranyume 1 hour ago
        I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.
        • guelo 1 hour ago
          Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.
          • ranyume 55 minutes ago
            But we do know they're "intelligent" and also smart in an important capacity. So how gives we don't measure them by IQ? Because the IQ is not a good measure of intelligence or smarts.
            • mdp2021 44 minutes ago
              > Because the IQ is

              No, it is just because they have difficulties at the bench.

              > how gives we don't measure them by

              We'd measure them by all the tests available. Not all test are usable in all circumstances.

              • ranyume 18 minutes ago
                > No, it is just because they have difficulties at the bench.

                I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments for good which is one way we define intelligence.

                On the other hand saying "the gifted person is highly intelligent/smart just not good at some things" really diminishes the other tasks, because they really are very difficult tasks but are not measured by an IQ test.

    • jawiggins 2 hours ago
      There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
      • eli 1 hour ago
        Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.
      • ranyume 1 hour ago
        Give away access to the model and go ask people from time to time if the model was of use to the person and if they were able to make the model work with them.
      • malfist 2 hours ago
        I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it
        • hyperhello 2 hours ago
          This is like arguing about whether a hot dog is a sandwich (of course it is) or whether the chicken or the egg was first (obviously the egg since all chickens come from eggs). Intelligence is just problem solving in the context of self-awareness. Machines don't have it and never will but they can simulate the process given inputs. You can argue whether humans and animals truly possess self-awareness and in what degree, but the definition of intelligence is as simple as the hot dog debate.
        • whattheheckheck 1 hour ago
          It was defined in Animal Intelligence by George John Ramones in 1882 as "intelligence is the capacity to do the right thing at the right time. It is the ability to respond to the opportunities and challenges presented by a context"
    • rcoveson 2 hours ago
      No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.
      • malfist 1 hour ago
        You forget the benchmark. The human subjects were told they were being timed. If you believe the lowest time is the primary metric you will absolutely trial and error at speed instead of meticulously plan out your moves to minimize that metric.

        LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.

        So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.

        This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs

        • mdp2021 51 minutes ago
          > LLMs are not timed

          Not fully relevant: timing is crucial in all-pass tests, not crucial in pass-or-fail tests. I.e.: first of all, they have to be able to reach the goal, and that is already an achievement. Then - and in parallel - the problem solving must also be optimized for efficiency. But "solving" and "efficiency" are non coincident dimensions.

    • jrflo 1 hour ago
      You should read more on the ARC prize, it actually has a pretty long history. We're on the 3rd iteration because they keep getting saturated. If you look at the score history over time on ARC AGI 1, 2 and 3 it's pretty impressive.

      https://arcprize.org/

    • paimapi 1 hour ago
      I'm also unclear as to how basic inferential logic puzzles spells out intelligence

      I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations

      [0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...

      • mdp2021 49 minutes ago
        > basic inferential logic puzzles spells out

        It spells out a form of intelligence - some can and some cannot.

        Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.

    • dist-epoch 1 hour ago
      1.5 years ago Gemini Pro 2.5 needed 1 page of thinking for every move in tic-tac-toe.

      Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.

    • GaggiX 2 hours ago
      If you have never seen the game before probably.
    • nimchimpsky 1 hour ago
      [dead]
  • fxd 20 minutes ago
    “AGI” never made sense to me. It’s a purely marketing term right?

    I’ve ignored it thinking it would go away, but it keeps coming up.

    I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.

    Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.

    But you’d be no nearer to solving consciousness.

    Given this thought trajectory - what is AGI supposed to be?

    • layer8 9 minutes ago
      Not sure why you are bringing up consciousness, that’s largely orthogonal to intelligence. AGI is usually taken to mean the capability to match or surpass human intelligence across all conceivable cognitive tasks, as opposed to being limited to certain kinds of tasks, or to not matching the general level of human intelligence in some respect.

      Intelligence, and hence AGI, doesn’t require consciousness or emotions or sentience.

      • fxd 3 minutes ago
        [dead]
    • eagerpace 14 minutes ago
      I like recursive self improvement instead. It seems like something that is actually quantifiable and kinda “the point” of why consciousness is important to humans.
      • fxd 5 minutes ago
        So basically, being able to set it free on some long running goal and it sort of “lives” and autonomously does its own tasks?

        I wonder at what point consciousness is necessary… that is, if you can have anything like that without it.

        To the point that solving consciousness (and combining it with intelligence) is what gives you the autonomous, recursive, self-improving thing otherwise it can only drive in the dark and make big mistakes.

        To your point I think - it’s why we don’t see too many non-conscious advanced biology (it rarely survives against those with it).

  • Betelbuddy 2 hours ago
    "For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

    Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

    Well I dont know about all of you, but I am celebrating meat based humans...

    • LPisGood 2 hours ago
      I think raw brain energy is not a fair comparison. Humans are not willing and able to serve requests at identical competence all hours of the day. You have to invest considerable resources to get a person to even do so for part of the day.
    • paxys 2 hours ago
      Why are you making the assumption that a person's time is worthless? I'd argue that it is the single most valuable resource we all have.
  • fastball 28 minutes ago
    Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?
  • 6thbit 1 hour ago
    The instant/no reasoning performed extremely well

        none 35.2%, $49,791 96.7%, $23,457
    
    35.2% on the standard harness, that's above Opus 5 on high.
    • NitpickLawyer 1 hour ago
      Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).
  • dwohnitmok 2 hours ago
    > Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.

    Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"

  • mikert89 2 hours ago
    Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
    • tedsanders 1 hour ago
      Disagree.

      Examples:

      - predict a coinflip: easy to verify, hard to learn

      - earn $100: easy to verify, hard to learn

      - increase paid subscriptions in an A/B test: easy to verify, hard to learn

      I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

      • ranyume 57 minutes ago
        Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.
      • mikert89 1 hour ago
        these just need more compute:

        - earn $100: easy to verify, hard to learn

        - increase paid subscriptions in an A/B test: easy to verify, hard to learn

        but we both know these examples go against the spirit of my point

        • tedsanders 1 hour ago
          Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).
          • mikert89 55 minutes ago
            yeah but I think you may be underestimating the amount of capital available for compute. if AGI is possible through some 5 trillion of expenditure on computers, there will be money for it.

            also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye

          • _superposition_ 53 minutes ago
            You bring up an interesting point. Isn't the reward itself subjective in many domains?
    • jdthedisciple 15 minutes ago
      Yes, but not necessarily under tight budget constraints.
    • x3haloed 2 hours ago
      Yup. Only subjective taste remains.
      • GPerson 1 hour ago
        Nope that will be commodified in short order.
  • yomismoaqui 26 minutes ago
    Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI?

    Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)

  • hypfer 2 hours ago
    What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?
    • petu 2 hours ago
      OpenAI provides API key with ~unlimited use?
      • Frost1x 1 hour ago
        So, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute.

        Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.

        Alignment++

  • piloto_ciego 2 hours ago
    99.9% with the right harness? Ok, we're at AGI then.

    Prediction:

    We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.

    • WASDx 1 hour ago
      They explain it here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

      TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".

    • jhonof 2 hours ago
      • NitpickLawyer 1 hour ago
        AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.
      • piloto_ciego 2 hours ago
        I rest my case.
    • emp17344 2 hours ago
      Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?
      • piloto_ciego 1 hour ago
        Didn’t I just see a thing about how actual unemployment is at like 24% a few days ago?
        • raspasov 1 hour ago
          According to that interpretation, ~24% is one of the lowest ever.

          https://www.lisep.org/tru

          (I have not gone down the rabbit hole to understand how they achieve that 24% number)

          • _superposition_ 50 minutes ago
            Workforce participation is different than unemployment. Didn't click the link but I suspect that's the case with your 24%
    • dgellow 2 hours ago
      The goal moving is by design, that’s why they use something as ill defined as AGI
    • slopinthebag 2 hours ago
      That’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.
      • piloto_ciego 2 hours ago
        I agree, like the average human isn't generally intelligent.

        IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!

    • baal80spam 2 hours ago
      > We will now see the goalposts moved

      It's already happening :)

      • piloto_ciego 2 hours ago
        Hilariously it is, I'm just reading more on this!
  • yusufozkan 3 hours ago
    what the hell is that score/cost curve lol
    • minimaxir 3 hours ago
      DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.
      • Frost1x 1 hour ago
        It’s not that different than a lot of real world economies. Often paying for someone or something with better quality can reduce total costs. You have less failures, less mistakes, so on, so while the expertise or quality of the product is higher than cheaper solutions, they can be more reliable and over time ultimately cheaper.

        The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).

      • Phemist 2 hours ago
        What is the intuition. Higher quality turns due to more reasoning results in significantly fewer turns taken?
        • tedsanders 1 hour ago
          Yep. In particular, ARC-AGI-3 is a series of games where if you fail, you keep trying again (until eventually hitting a timeout). So the sooner you succeed, the sooner you stop spending tokens retrying. If it was a benchmark where everyone got one attempt with no retries, you wouldn't see it bend backward.
        • minimaxir 1 hour ago
          Yes, in theory.
  • bigbuppo 1 hour ago
    Wake me when it's going to spontaneously fix my leaky faucet because if it doesn't do it nobody else will. Until it has that capability I don't really care.