Ember-1

(fireworks.ai)

122 points | by gmays 1 hour ago

18 comments

  • GodelNumbering 24 minutes ago
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
    • shriphani 17 minutes ago
      what hardware are you using to train?
      • GodelNumbering 13 minutes ago
        I didn't have a local GPU, so I asked it to go out and find hardware. It found a google TPU v6e which seemed reasonably priced. I gave it my google api key. I told it to use TPU only when training and bring it down afterwards. That's about it.
        • otterley 5 minutes ago
          What kind of observability did you have over this process? I’m interested in how my peers are operating these efforts.
        • shriphani 9 minutes ago
          Neat!
    • PEe9bB7D 15 minutes ago
      i also need more info!
      • GodelNumbering 12 minutes ago
        I am thinking about opensourcing everything, although this is not my main domain or my main startup, so the overhead of huggingface etc seems a bit unnecessary
  • netvarun 52 minutes ago
    Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15
    • drob518 38 minutes ago
      Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost.
    • nostrebored 30 minutes ago
      Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design
  • jamienk 1 hour ago
    Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
    • andsoitis 51 minutes ago
      > Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

      I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.

      So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.

      • jamienk 44 minutes ago
        Diff people have diff motives to experiment, then new work is done on top of stuff that "hits" in a way no one anticipated. Then work gets piled on top in a way that might make it hard to port
        • andsoitis 43 minutes ago
          > then new work is done on top of stuff that "hits" in a way no one anticipated.

          Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.

          Greatness cannot be planned.

    • segmondy 51 minutes ago
      No, because close labs/models borrow but don't contribute back.
    • swagatkonchada 26 minutes ago
      Won't the "frontier" labs figure out whatever techniques were used and apply them to their closed models?
      • cyanydeez 16 minutes ago
        Like how the last 2 decades of tech companies are thinly veiled open source pilfering into business units.
      • k__ 19 minutes ago
        If they can keep up.

        The lock-in is less pronounced as it is with AWS or MS.

    • zeroq 28 minutes ago
      The difference between contributing to OS and AI, is that the first is a hobby alternative to woodworking or hiking, while the other can easily bootstrap you a company you can get millions in investment, at least for time being.
    • intothemild 58 minutes ago
      Yes, absolutely, but only if people keep contributing in the open.
      • jack_pp 51 minutes ago
        not necessarily, just knowing something is possible will motivate others to achieve it somehow. Which is why there are so many LLMs and OAI doesn't have a monopoly
  • nico 40 minutes ago
    > The problem: thinking models think too much

    This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking

    It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications

    • elcomet 12 minutes ago
      Why not using a cheap LLM with thinking completely disabled ? I don't think it will be much more expensive than jev.
    • demibabs 30 minutes ago
      What are the useful applications of Jev so far? Not to sound dismissive, I just haven’t seen what people are using it for yet.
      • neosat 25 minutes ago
        Lots of use cases! I've personally used it for the following:

        1. Evals (once you have your rubric defined and tuned using a reasoning model, jev can be great for running periodic evals especially those that run daily.

        2. e-commerce catalog classification 3. quick search using anything as context and query mapping to a pre-defined set.

  • andsoitis 1 hour ago
    > The problem: thinking models think too much

    Analysis paralysis stifles not just human intelligence, but other intelligences too.

    • AraneaDev 4 minutes ago
      Yes and thar makes you wonder if the Paradox of Choice would apply as well ;)

      The more options you have, the harder it becomes to be satisfied with the one you picked.

  • intothemild 59 minutes ago
    So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?
    • DonsDiscountGas 39 minutes ago
      It happens. Most open licenses aren't GPL style copyleft.
    • netvarun 49 minutes ago
      Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.
      • Evidlo 38 minutes ago
        weight available
    • makeramen 26 minutes ago
      Aren't Cursor Composer models like this too? At some point all the extra RL you do can be considered as proprietary information added.

      Not suggesting this is right or wrong, but is sort of the nature of the technology.

    • kingstnap 23 minutes ago
      There is little to no point reading the article as well. It's stripped of all alpha.

      > task and environment feedback

      > on-policy planning and learning

      > feedback connects decisions to their consequences

      These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.

      Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.

    • swagatkonchada 29 minutes ago
      It happens with open source software all the time, why would we expect any different with open source weights.
      • otterley 4 minutes ago
        Because the licenses that apply to software make no sense in the context of LLMs. With the latter, there is no source code to license.

        The words of a license are what the license is.

      • reactordev 26 minutes ago
        Because we do. The GPL isn't a suggestion. If you can take open source code and make private software out of it then what are we all doing? No, license requirements and agreement are law for a reason.
  • tomrod 1 hour ago
    Well done, and great iteration.

    The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).

    • drob518 37 minutes ago
      Unfortunately, it’s hard to make a chart of that.
  • tdhz77 1 hour ago
    Does anybody know if this would be a good model for creative writing?
  • erichocean 1 hour ago
    Need this done for DeepSeek, ideally one of the Flash models.
    • drob518 35 minutes ago
      And GLM. Both Deepseek 4.1 Flash and GLM 5.3 Flash are quote verbose when thinking.
    • atemerev 1 hour ago
      If you have the compute, I have the expertise.
  • dbuxton 41 minutes ago
    Do they mean Opus 5.5 or Opus 5?
  • themgt 40 minutes ago
    The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.

    "Pareto": 8 hits

    "Opus 5.5": zero hits

    • wmf 32 minutes ago
      Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.
  • ls612 1 hour ago
    On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
  • monkey_monkey 1 hour ago
    I don't think the article mentions Pareto frontier enough.

    Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?

    • user43928 24 minutes ago
      Pareto frontier on some benchmark that I am hearing of for the first time.

      Kimi K3 with less reasoning tokens isn't exactly exciting either, and particularly so if the license is less open than original Kimi K3.

    • AnodicElegy 52 minutes ago
      I guess they figure "best bang for your buck" comes off a little too colloquial.
    • DonsDiscountGas 36 minutes ago
      They want it to be the best at something. And it's obviously not the absolute smartest. So here we are.
  • logicallee 43 minutes ago
    This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.
    • nostrebored 28 minutes ago
      I can’t help but think it’s more expensive tinker.
  • esafak 1 hour ago
    It looks like it would be similar to GLM 5.3 Flash, had they tested it...
  • fr2029 10 minutes ago
    [dead]
  • justmeeew 43 minutes ago
    [dead]
  • huflungdung 1 hour ago
    [dead]