7 comments

  • DiabloD3 27 minutes ago
    The article doesn't really describe the problem: if your prompt is 35kb, your prompt is confusing, unfocused, and doesn't work right on any LLM, and is needlessly bloating your context.

    At this point in time, due to how most people and companies run their inference engine, regardless of the model (yes, this includes the newest from OpenAI and Anthropic and the Chinese Tigers and Dragons), you run out of useful context that the model can accurately attend to around the 250k mark no matter how much they advertise their context size is.

    You need to cut your prompt up. If you believe LLMs work, have the LLM help you shape the overall plan, and then have multiple sessions run each step in the plan without being bloated with the context of previous successful steps.

    I don't see LLMs being production-ready until the context rot and sampling problem is fixed forever. This has not occurred, and the big inference providers aren't even bothering to integrate any of the research on that subject.

    If anything, many of the bigger companies are actively making inference quality worse just to extend their runway a tiny bit farther before they go bankrupt.

    The only thing the article gets right is this: if you're serious about LLMs, abandon Big AI and infer locally only. This is the only way you have control over the quality of the output.

    • andai 18 minutes ago
      The Claude Code system prompt was >50KB, though I think they trimmed it down heavily recently. (The newer models don't need as much hand-holding.)
      • embedding-shape 12 minutes ago
        Yeah, strong evidence for what parent says is correct. Been my experience as well, especially with local (smaller) models but also SOTA. The less instructions you have, the better they get at following them. Conflicting instructions is like poison, and it's harder to find those conflicting parts the longer the prompt is too.
  • cube00 5 minutes ago
    Friends Don't Let Friends Use Ollama https://news.ycombinator.com/item?id=47788385
  • andai 23 minutes ago
    > Everyone who begins learning exploitation hits a phase of exploitability grief about 3 month into dedicated, practiced study. They hack something they didn’t think they had the skill to break into and it terrifies them. They’re smart enough to know that, relatively speaking, they are an idiot, and if an idiot can do this then nothing is safe. That feeling is correct.
  • stackedinserter 16 minutes ago
    The main gotcha for local models is insane hardware requirements.

    Even for $10K you get mediocre performance.

  • SyneRyder 37 minutes ago
    TLDR: Local models have a smaller context window, so your 35kB prompts that worked fine against a hosted 1 Million token window, crash out when you only have a 65K (!) token window locally.

    I dislike being negative, but I was really hoping for more substance when reading this. It would have been an interesting topic.

    • 0o_MrPatrick_o0 30 minutes ago
      Thanks for the feedback. I wanted to get into more detail, but I spent the whole weekend working these problems and then constructing this post.

      Dario’s behavior this weekend made me feel like this just needed to get out quick. In the future, I’ll be sharing more details about some other things in the process and some ways I found to use automation to accelerate splitting prompts for use on local inference.

      • SyneRyder 14 minutes ago
        Understood, and I realized you'd posted this to HN yourself, so I felt a bit bad making the comment. I think maybe for me, this might have worked better if the motivation had been one separate post, and the details of the gotchas as a post of its own.

        But I also think, if you're going to be limited to 65k token windows, you're going to have a really difficult time. Even 250k windows were cramped for me when that's all we had on Anthropic models. I just don't think a 65k window is going to be big enough for proper cyberdefence work, even if I totally agree with going local wherever you can. It feels like if you're defending against swarms of 1-10M context windows, you need to get as close as you can to similar. I've had to reach for Chinese 1m models instead because the American models just refuse me here in Australia.

      • monegator 29 minutes ago
        Yes, please!
  • dell2024 28 minutes ago
    I had hoped to get some new information out of this topic, but unfortunately found the same local "dead-ends" that I explored myself.

    It unfortunately feels like we will be stuck waiting for a burst bubble before local hardware can be reasonably acquired for personal LLM usage.