I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.
It didn't copy any source code from any external projects. I had it write a stratified sampling renderer for ground truth, then had it implement feature by feature by matching the pixels. Unless you mean it "copied" it in the sense of third-party code being part of the training data. I don't think that definition of "copy" makes any sense given how these models represent embeddings. It would also imply that humans are "copying" the things they've learned from.
Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
You're using quantity metrics to answer a quality question.
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
If you really want to check some quantity metrics to try to reason about code quality, look at whether "fix" PRs are overall LOC neutral or negative (not counting tests). In this project, almost every "fix" is an addition. Worse, almost every fix is more branching.
If almost every PR is some sort of fix, and most of them add branching, and there's thousands of them weekly... That leads to only one place and I want to be nowhere near it.
I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
People, especially those who don't do software development, often conflate coding and software development.
The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.
That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.
>The models aren't good at architecture and design.
Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?
Because I found the current SOTA AI models being great at architecture and design, much better in fact than most average real-world devs. Is your yardstick just the John Carmacks of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann, etc.
Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?
Cool. Can you elaborate at which types of task you are better than SOTA LLMs in context of "being good at SW design and architecture"? Got any examples? Is it at the interview questions? Or real world problems? If so what is the scale of the real world SW design and architecture problems you're better than the SOTA LLMs? Is it FAANG scale or mom and pop shop scale?
And do you consider yourself to be representative of the average developer, above them, or below them?
LLMs don't even need to be better than the average dev, let alone the top performing ones like you. If they can be better than the bottom 20% of devs and white collar workers in general(easily achievable when you've been around the block and see how many useless people just keep warm chairs for high wages in large companies), that's already a huge win for those products.
I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.
Were you reading through the code it was generating to make sure the flow was intuitive and comprehensible for each PR? Projects only turn in to maintainable messes if you blindly merge in unmaintainable messy code.
As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
I treat AI like an intern or junior team member. With enough guidance, they can contribute a lot but you can't let them loose without supervision or they usually will produce a big mess. As of now, you are still responsible for overall architecture. One strong indicator that something is going wrong are big pull requests where the AI has rewritten large sections of the code.
Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023.
I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.
If you specify a structure, pattern, or design, it will stick to that design after a code review phase. Just tell it what standards you have, and after two rounds, it will have written what you described.
AI is a useful tool but it absolutely needs a lot of human guidance.
Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.
Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.
Vibe coding has always had and always will have one fundamental flaw: you didn't write the code it produced, therefore you don't have a good understanding/mental model of it.
On the other hand asking these clankers "review the feature branch I wrote" and "review my entire codebase for bugs" or "help me debug this" has saved me months of prospective work.
And more recently most major models have been getting _really_ good at RE, for example you can have OAI models (and maybe A/'s if they don't refuse) use idalib MCP and reverse-engineer stuff from start to finish, then follow up with GLM 5.3 for vuln assessment and exploit PoC.
Stuff that used to take weeks or months now just takes a few hours, or less.
Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).
Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.
If dev approaches zero, but you review at the same pace as you always have, are you in a better position? Yes.
Will you potentially have a backlog of code waiting for review? Also yes.
Would you prefer to be waiting for the dev team for all of the time instead, then still have the same amount of reviewing to do at the end of it? Absolutely not.
It's the same as any other PR. Keep the changes small and contained to that feature. AI can do this if you instruct it to, there is no need to vibe code some 40k line monstrosity.
I am not saying you are right or wrong, but I don’t think you can make such a broad conclusion simply because of your own experience.
You say you used AI and your projects turned into unmaintainable messes, so your conclusion is that it means AI is not living up to its promises.
I guess if the argument is “AI makes it so you always get a great result no matter how you use it”, then your argument is sound. Your projects not working out proves that AI doesn’t always work no matter what.
However, that doesn’t mean you can’t use AI to create sustainable and well organized code. Failing to do something doesn’t mean it’s impossible and anyone who thinks they can is not paying attention.
I can’t run a marathon. If I went out and tried to run one, I would get a few miles and collapse, failing completely.
I don’t think it would be reasonable, though, at that point to say “running a marathon is impossible, anyone who says they can do it clearly lying. I tried and didn’t even make it 5 miles!”
I wish people would stop assuming their experience with something is the only possible truth.
The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.
I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.
I think it depends which LLM tool you're using. If you're using an older, worse model (the kind that you can use for free), the experience is significantly more frustrating. On the other hand, I'd say that the current best models are very useful with a skilled operator.
There are a couple of notable counterexamples here (nobody sane thinks Carmack doesn't know how to program, for example), but by and large I agree with your observation. The people excited about programming with LLMs are, on average, people who weren't good at programming to begin with. Still, given that these counterexamples do exist I try to avoid painting with an overly broad brush for the sake of nuance.
You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?
It also doesn't answer the question of how they might even recognize LLM generated code in contributions to PopOS directly.
I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.
I daily drove the alphas before the betas, and of course there were a couple rough edges. But I had a minimal working desktop instead of sway or KDE/Gnome (too heavy).
It has been releasing non-betas for a good year now, and it's been a smooth sailing.
This is a good idea, we should start a blacklist of open-source projects that are known to have used LLMs. There should be two universes of code, one for hand-typed code used by people who care about quality and one for slop used by those making trash.
Eh, not really. It just lowers the barrier for lazy people who never would have tried contributing before.
I’ve read a lot of code hand-written by programmers. By smart and hard working people. And I know from experience that the code the average programmer writes is not great either.
Yeah but that's just bad code in general, no? You can make good code with LLMs, you just have to actually engineer it and give up some of the velocity; which is just a bigger version of the same problem we've always had (yes, I get that code review can't scale).
This is just really silly.
The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.
I had a PR in flight that got closed because of this. I had an issue with passwords in the network applet for the VPN and had used Claude to help me identify and then come up with a fix. I did spend a lot time handcrafting and making sure the quality was good, but I respect their decision and no hard feelings, but as someone who have struggled to find time and opportunity to contribute to open source it was a small set back.
I found your commit and your usage of AI seemed reasonable. It seems to me like your PR itself and the subsequent comments and correspondence was also human written.
I think a PR "in flight" shouldn't have been closed like that.
All this will do is push out developers like you that honestly disclose, and instead people will now just lie.
What I meant by it was that I read every line of code generated, made sure I understood its purpose and either manually rewrote it if I felt there was a better way or asked Claude to do it. E.g. there was a bug where the password could end up in a configuration labelled as a username. Claude's initial fix was to simply exclude 'username' in a for loop, but I asked it to find examples in similar code in other codebases to see what the best practice was and ended up basing to fix on what is done in Gnome.
My guess: taking personal responsibility for the functionality, readability, and sanity of the change proposed, both atomically and in the context of the wider code base (adhering to existing conventions and patterns), to the best of the author’s ability.
Putting any specific string in an "author(s)" or "committer" field is likewise a token gesture that doesn't affect or alter the content. The impact, just like promises of responsibility, would still solely be in biasing the reader, clouding the evaluation.
Unless you believe there is substance in social interactions over time, like trust. But then it would be a socially weird move to dismiss promises of responsibility out of hand instead of picking up the invitation to build trust if that's what's perceived to be lacking.
I don't see how this will survive the attacker/defender gap as ls get increasingly good at cyber security and finding 0 days... but maybe it's an obscure enough is it doesn't matter?
Probably because pop_os and Cosmic are so niche and their market share so insignificant, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry to make their time and effort worth it. See the Arch AUR attacks, for perspective.
I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox, accepting only human written code.
But larger and more important projects like Fedora and Debian are more pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.
The thing is, the cat's out of the bag on this one now, especially in the field of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat even experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.
You only need one bad actor. For example, someone reading this thread could easily decide to start attacking it just because someone else said it wasn't worth it, as a personal challenge.
If that's your threat model then you shouldn't use any SW in exitance, FOSS or otherwise. In fact you shouldn't even go online, or even outside you own house, since one single bad actors exist everywhere. You can walk down the street and suddenly someone in a car runs you over(witnessed myself). And yet live goes on.
It doesn't have to be niche (or not) to have bad actors. That doesn't make it likely of course but it is possible, and increasingly more so in the age of AI when you can simply point it at multiple projects in parallel.
I never said they should. I think you're reading too much into my comment, the point was something being niche doesn't prevent it from being threatened.
> Probably because pop_os and Cosmic are so niche and their market share so insignificant
Consider that it's packaged for many well-known distros, so pop_os install base alone doesn't tell the whole story: https://system76.com/cosmic/download
So then COSMIC isn't even TOP 5, for this to be a major target by market share as originally claimed.
2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.
3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.
To reflect some sentiment from another comment you posted - we're not your unpaid tone auditors. Learn to have normal discussions.
For funsies, I threw two different DE-by-marketshare inquiries at Gemini 3.8 Pro in separate sessions and it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.
Please show me what part of what I said before was not "a normal discussions" according to you?
>it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.
And that disproves me how exactly? I originally showed "Cosmic is NOT a TOP5 DE", and the LLM data you posted also shows that IT IS indeed NOT a top 5 DE.
"The analytics", is the publicly available training dataset of the LLMs that share the same common opinion on that DE market share.
What do you expect exactly? Do you want me to now manually parse through terabytes of information at your whim for your own convenience? Sorry, but I'm not your personal unpaid servant.
If you wish to disprove me in the comment section, then you need to do the manual work and show us that the my quoted LLMs statistics are wrong. I'm not your personal errand boy to do your bidding, 'massa'.
Point at the public dataset then if it's so public. You really have no idea how the LLM is parsing the data and could easily be hallucinating. That you find being called out on this as some sort of insult is quite telling and frankly funny, what a strong reaction to someone asking for a basic source when the burden of proof is on you to prove (or at least show a source that) these stats are right, not on anyone else to disprove AI bullshit.
Sorry, it's not my job to provide for you the things that you demand from me at your whims, because the original comment I replied to with LLM data also did not bring peer-reviewable information as proof that COSMIC is a TOP 5 DE, and yet you did not demand proof from them in that case. Why is that?
Did you just blindly trust the opinions of others on this topic, and yet I'm the one who has to provide peer-reviewable data for you to back-up mine? Sorry, I'm not your unpaid lackey. Try to formulate a better (counter)argument for why their baseless opinion is right, but my LLM backed up opinion is wrong, if you wish for a even-footed good-faith argument.
I ask them for the same source (and looks like they provided it), your comment wasn't special. LLMs however are especially less trustworthy, that's why. It's not my job to educate you on why they are, as you say, and why people aren't trusting your comments.
>I ask them for the same source (and looks like they provided it)
I also saw it now. That blog is not a representative ground truth, but just another opinion piece, which I can respect as an opinion of the blog's user base, but I can't take as an accurate real world statistic, same how aggregate opinions you read on HN are not representative of the actual real world.
I hope you can understand my PoV. You can also disagree if you want, but you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.
In theory if you brought a better source than that person then I'd agree with you, but,
> you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.
actually I can say it's wrong or likely to be wrong especially if it doesn't cite the sources it uses. And if it does, then just paste the sources here instead of the LLM output. It is also unknown where it got the info and as someone else said, you ask it two different times and it gave two different answers, thus it is unreliable.
OK, but what's the sample size of that and who's measuring it? I never heard of that blog or took part in that poll. So how is that blog link the yardstick but mine is not?
>Because that blog actually talked to real people and did real wor
How did you verify that those people from the blog are "real"?
I also talked to real people for my own data, case in point, I asked my mom and dad which linux DE is most used, and the results came out different. Which "real people" are the ones representative for the ground truth of Linux DE sahre?
> and yours is just a hallucinated list from one of the me-too LLM vendors
How do you know it's hallucinated? Ask the LLM the population of your country? Is the answer mostly accurate or is it hallucinated in an inaccurate way?
Aren't LLMs just outputting the highest statistical probability from the aggregate of their scraped data, which in this case would be including opinions on Reddit, and every website and blog on the entire internet (including that random one posted by badc0ffee) on the Linux DE uusage topic, making it a more accurate real-world representation than just a single random blog?
You can call it "hallucinated" if you want, but that doesn't mean it's not accurate. I asked for proof that my answer was inaccurate, not that it was "hallucinated", those are two different things, and your argument didn't prove it was inaccurate nor did it prove it was hallucinated. Would you like to try again?
<< Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.
That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.
<< I think even amongst the HN and Linux userbase, pop_os is still niche
I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?
It also means security is not held as high and vulnerabilities not as much found. A simple 0-day may survive for years. Not much effort needed to have permanent access.
Its easier for an LLM to find vulnerabilities in a project with less usage, and those vulnerabilities will stay open longer, making them much cheaper to attack to keep the door open.
What's the point of attacking projects that almost nobody uses?
Do you think Netanyahu, Trump or Xi-Jinping are somehow secretly using Cosmic DE at home, to be worthy targets?
Bad actors have limited time, lives of their own and mouths to feed as well, so they concentrate their efforts where "the fish are" if they want to PWN someone for profit.
That's why Windows was the biggest target in the past for so long and why MacOS and Linux were ignored. Because most of the fish were on Windows.
Offensive security employees, tokens, and peoples' time are still a finite resource that get allocated based on target priorities and operational end-goals, even by state actors.
If you assume Mossad and NSA are Token-maxxing every single niche FOSS project out there to cast as large as possible fishnet on hacking all Average Joes on the planet just in case, then maybe using Mozilla and MacOS gets you hacked too, maybe even visiting HN and commenting here gets you hacked by some zero days you don't yet know.
So you took every single line of open source code you could possibly get your hands on (using scrapers so violently dumb that they amount to a permanent low-grade DDoS) and spent billions of dollars to tune trillions of parameters, and the value you can offer is… “let us inundate you with bad code or else we’ll generate exploits for your software”.
Using AI to find vulnerabilities doesn’t mean that you need to use AI to generate the code that fixes them. And you can still ask AI whether it thinks the fix is okay, as a second opinion.
I wonder if the issue is mostly the code or the AI written PRs and people using AI to talk to the maintainers. I personally just ban anyone doing the latter, I don't want to talk to opus more than I already do lol
Issue is entrenchment in the old SWE world that ceased to exist somewhere in August this year, and using AI to mechanize the traditional workflows. Also people with "I need to eyeball each char in PR diff" attitude, which is fucking unproductive at this point.
Ok, this may be controversial, but LLM code tokens aren't free, and I run out of my weekly allowance pretty regularly just from doing some fairly heavy projects, so I don't understand why somebody would ever want to spend their own money to make bad PRs on purpose, and I like to assume good intentions from people unless proven otherwise, which means a near blanket ban for LLM authored code for these big open source projects just seemed a bit extreme to me, when the core issue seemed to be that review process/policy should change with the times.
For example, I was helping work on an open-source game engine earlier this year with a longstanding text rendering bug dating back to around 2021 that prevents the engine from being production ready, which the community and myself have developed extensive workaround for. So, one day I've finally said enough and got Claude to debug it. It took Claude 10 minutes to find the bug, it was 3 lines of code change in the renderer (yes, three).
So, I wrote up the regression tests, documented the bug and opened up a PR for the fix, thinking it'll get merged in like less than a week and then we can all move on. The maintainers received it fairly well on the PR, but the PR sat there for nearly 6 months, unmerged, until it finally closed from a bad squash upstream. I'm pretty sure the bug is still there too.
And as a side note, I would be ecstatic if someone wants to contribute to my Github projects with their AI.
"... many of the AI contributions were not planned and showed little understanding of the software architecture. So the team wants to "prioritize working on contributions from our own team and regular contributors.""
Sounds reasonable, even to avid LLM users, I suppose. You have to draw a line. This line is too simplistic, but it'll work, for now.
This would just push people to maintain their own fork. If I already have an agent to investigate and fix a bug and able to send a PR, the added cost of maintaining a local fork is minimum.
In fact I've start doing that myself. Sending PR and convincing the maintainer why the fix is necessary is just too much effort.
Same trouble we have. Some clever person says to use AI agents for code review. 100kloc commit got flagged through on Friday. Taking this week off. Not my circus.
Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.
There are a whole lot of people (in tech) who truly hate AI and want nothing to do with it. Those people will flock to projects who take a stand against it.
it doesn't matter. it's delusional to think you can outcompete a thing for which solving a Millenium problem is just Tuesday. it's the anger phase of grief, nothing more.
This is straight up cult behavior. "You WILL assimilate or we WILL kill your project"
PopOS is not a product being sold by techbros trying to pump their stock like AI is, it's free open source software, it doesn't need to "outcompete" anything. If you don't like it, don't use it.
it isn't even me not liking it (I don't think about it at all frankly, I have nothing to like or dislike) - I think it's irresponsible policy. I won't be using something that doesn't accept automated security patches from defender AIs.
I won't and I recommend anyone currently using it to reconsider due to the security posture the project necessarily and unfortunately has as a direct consequence of its policies.
banning AI from PRs because you're swamped with too many low quality PRs, definitely. We pretty much are doing this with SQLAlchemy. If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Reasonable, although I've taken a different approach. Either closing such PRs, or treating them as very detailed issues and having my own LLM build the actual fix.
My repos probably don't see as much traffic as SQLAlchemy though.
rejecting low quality ones should be the norm regardless of whether an AI or a human wrote them. the question is what would happen if you were swamped with high quality PRs? what will happen once you are? (that's probably a 2027 question!)
I'd still reject them. It's the bug reports and feature requests that are valuable. When you have an AI yourself, there is little point in having somebody else let their AI implement them, that just creates a lot of risks and unknowns for no benefit.
> If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
I’m saying it’s probably multiple factors and both you and GP are right.
A good workman shuts up and finds better tools without complaining.
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
There's a blog entry https://nousresearch.com/refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.
I guess they know how to hold it?
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
If you really want to check some quantity metrics to try to reason about code quality, look at whether "fix" PRs are overall LOC neutral or negative (not counting tests). In this project, almost every "fix" is an addition. Worse, almost every fix is more branching.
If almost every PR is some sort of fix, and most of them add branching, and there's thousands of them weekly... That leads to only one place and I want to be nowhere near it.
Show me an AI that adds features by deleting code (https://www.folklore.org/Negative_2000_Lines_Of_Code.html) and I'll pay attention.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
What was your process?
The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.
That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.
Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?
Because I found the current SOTA AI models being great at architecture and design, much better in fact than most average real-world devs. Is your yardstick just the John Carmacks of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann, etc.
Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?
And do you consider yourself to be representative of the average developer, above them, or below them?
LLMs don't even need to be better than the average dev, let alone the top performing ones like you. If they can be better than the bottom 20% of devs and white collar workers in general(easily achievable when you've been around the block and see how many useless people just keep warm chairs for high wages in large companies), that's already a huge win for those products.
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.
Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.
On the other hand asking these clankers "review the feature branch I wrote" and "review my entire codebase for bugs" or "help me debug this" has saved me months of prospective work.
And more recently most major models have been getting _really_ good at RE, for example you can have OAI models (and maybe A/'s if they don't refuse) use idalib MCP and reverse-engineer stuff from start to finish, then follow up with GLM 5.3 for vuln assessment and exploit PoC.
Stuff that used to take weeks or months now just takes a few hours, or less.
`total = dev + review`
If dev approaches zero, but you review at the same pace as you always have, are you in a better position? Yes.
Will you potentially have a backlog of code waiting for review? Also yes.
Would you prefer to be waiting for the dev team for all of the time instead, then still have the same amount of reviewing to do at the end of it? Absolutely not.
You say you used AI and your projects turned into unmaintainable messes, so your conclusion is that it means AI is not living up to its promises.
I guess if the argument is “AI makes it so you always get a great result no matter how you use it”, then your argument is sound. Your projects not working out proves that AI doesn’t always work no matter what.
However, that doesn’t mean you can’t use AI to create sustainable and well organized code. Failing to do something doesn’t mean it’s impossible and anyone who thinks they can is not paying attention.
I can’t run a marathon. If I went out and tried to run one, I would get a few miles and collapse, failing completely.
I don’t think it would be reasonable, though, at that point to say “running a marathon is impossible, anyone who says they can do it clearly lying. I tried and didn’t even make it 5 miles!”
I wish people would stop assuming their experience with something is the only possible truth.
As ai becomes better these people will begin changing their story because it’s utterly obvious what’s happening.
You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?
I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.
PopOS is Ubuntu with extra problems. Ubuntu itself is fine, but then PopOS adds weirdness.
Cosmic has been in beta for how long ?
I daily drove the alphas before the betas, and of course there were a couple rough edges. But I had a minimal working desktop instead of sway or KDE/Gnome (too heavy).
It has been releasing non-betas for a good year now, and it's been a smooth sailing.
They aren't saying that AI produces bad code or is terrible for the world in some way.
It's mainly just resulting in a lot of PRs that they don't have enough time to review or features they don't plan to add.
Being hand written is no guarantee of high quality, just like using LLMs is no guarantee of low quality.
I’ve read a lot of code hand-written by programmers. By smart and hard working people. And I know from experience that the code the average programmer writes is not great either.
This is just really silly.
The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.
I think a PR "in flight" shouldn't have been closed like that.
All this will do is push out developers like you that honestly disclose, and instead people will now just lie.
Unless you believe there is substance in social interactions over time, like trust. But then it would be a socially weird move to dismiss promises of responsibility out of hand instead of picking up the invitation to build trust if that's what's perceived to be lacking.
I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox, accepting only human written code.
But larger and more important projects like Fedora and Debian are more pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.
The thing is, the cat's out of the bag on this one now, especially in the field of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat even experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.
If that's your threat model then you shouldn't use any SW in exitance, FOSS or otherwise. In fact you shouldn't even go online, or even outside you own house, since one single bad actors exist everywhere. You can walk down the street and suddenly someone in a car runs you over(witnessed myself). And yet live goes on.
Because your comment didn't disprove that Cosmic DE isn't too niche for bad actors to get involved.
Then they shouldn't reject AI aids to help them patch vulns found by bad actors faster, no?
Consider that it's packaged for many well-known distros, so pop_os install base alone doesn't tell the whole story: https://system76.com/cosmic/download
(I wasn't able to make it work on a scrap Dell I tried it on because the GPU was too old. Booted the USB key and COSMIC greeter failed to start)
Edit: I see you posted elsewhere.
1) Firstly, your source plase? My research according to Google Gemini 3.8 Pro shows top 5 DEs are as follows:
So then COSMIC isn't even TOP 5, for this to be a major target by market share as originally claimed.2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.
3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.
"Way too defensive" how? By asking and bringing data for my PoV?
>GP replying to you was clearly trying to have a conversation
As am I, except I ask for, and also bring data to back up my PoV, instead of vague opinions.
>Chill.
Where am I not being chill?
For funsies, I threw two different DE-by-marketshare inquiries at Gemini 3.8 Pro in separate sessions and it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.
Please show me what part of what I said before was not "a normal discussions" according to you?
>it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.
And that disproves me how exactly? I originally showed "Cosmic is NOT a TOP5 DE", and the LLM data you posted also shows that IT IS indeed NOT a top 5 DE.
Do you have a better source than GP?
>if needed post the actual source
Define "actual source"? In good faith, I mean.
Where else do you get this information that's, quote, "actual source"?
What do you expect exactly? Do you want me to now manually parse through terabytes of information at your whim for your own convenience? Sorry, but I'm not your personal unpaid servant.
If you wish to disprove me in the comment section, then you need to do the manual work and show us that the my quoted LLMs statistics are wrong. I'm not your personal errand boy to do your bidding, 'massa'.
Did you just blindly trust the opinions of others on this topic, and yet I'm the one who has to provide peer-reviewable data for you to back-up mine? Sorry, I'm not your unpaid lackey. Try to formulate a better (counter)argument for why their baseless opinion is right, but my LLM backed up opinion is wrong, if you wish for a even-footed good-faith argument.
I also saw it now. That blog is not a representative ground truth, but just another opinion piece, which I can respect as an opinion of the blog's user base, but I can't take as an accurate real world statistic, same how aggregate opinions you read on HN are not representative of the actual real world.
I hope you can understand my PoV. You can also disagree if you want, but you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.
> you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.
actually I can say it's wrong or likely to be wrong especially if it doesn't cite the sources it uses. And if it does, then just paste the sources here instead of the LLM output. It is also unknown where it got the info and as someone else said, you ask it two different times and it gave two different answers, thus it is unreliable.
How did you verify that those people from the blog are "real"?
I also talked to real people for my own data, case in point, I asked my mom and dad which linux DE is most used, and the results came out different. Which "real people" are the ones representative for the ground truth of Linux DE sahre?
> and yours is just a hallucinated list from one of the me-too LLM vendors
How do you know it's hallucinated? Ask the LLM the population of your country? Is the answer mostly accurate or is it hallucinated in an inaccurate way?
Aren't LLMs just outputting the highest statistical probability from the aggregate of their scraped data, which in this case would be including opinions on Reddit, and every website and blog on the entire internet (including that random one posted by badc0ffee) on the Linux DE uusage topic, making it a more accurate real-world representation than just a single random blog?
You can call it "hallucinated" if you want, but that doesn't mean it's not accurate. I asked for proof that my answer was inaccurate, not that it was "hallucinated", those are two different things, and your argument didn't prove it was inaccurate nor did it prove it was hallucinated. Would you like to try again?
That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.
<< I think even amongst the HN and Linux userbase, pop_os is still niche
I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?
Do you think Netanyahu, Trump or Xi-Jinping are somehow secretly using Cosmic DE at home, to be worthy targets?
Bad actors have limited time, lives of their own and mouths to feed as well, so they concentrate their efforts where "the fish are" if they want to PWN someone for profit.
That's why Windows was the biggest target in the past for so long and why MacOS and Linux were ignored. Because most of the fish were on Windows.
Previously, time was the most precious resource. Now its tokens, and more cheaply at at.
If you assume Mossad and NSA are Token-maxxing every single niche FOSS project out there to cast as large as possible fishnet on hacking all Average Joes on the planet just in case, then maybe using Mozilla and MacOS gets you hacked too, maybe even visiting HN and commenting here gets you hacked by some zero days you don't yet know.
Where does this open-ended paranoia argument end?
Reminder that AI is quite stupid.
I'll trust the words of groups like curl (https://daniel.haxx.se/blog/2026/06/10/a-human-in-control/), Linux, and even the infamously anti-AI Gnome (https://blogs.gnome.org/mcatanzaro/2026/10/02/the-era-of-sof...) that AI is finding real vulnerabilities and you're your project a disservice by ignoring them.
Edit: Though Greg did recently have a talk (that I skimmed) where he was a little reserved on LLMs: https://www.youtube.com/watch?v=NnV_cWeoo5Q
For example, I was helping work on an open-source game engine earlier this year with a longstanding text rendering bug dating back to around 2021 that prevents the engine from being production ready, which the community and myself have developed extensive workaround for. So, one day I've finally said enough and got Claude to debug it. It took Claude 10 minutes to find the bug, it was 3 lines of code change in the renderer (yes, three).
So, I wrote up the regression tests, documented the bug and opened up a PR for the fix, thinking it'll get merged in like less than a week and then we can all move on. The maintainers received it fairly well on the PR, but the PR sat there for nearly 6 months, unmerged, until it finally closed from a bad squash upstream. I'm pretty sure the bug is still there too.
And as a side note, I would be ecstatic if someone wants to contribute to my Github projects with their AI.
Sounds reasonable, even to avid LLM users, I suppose. You have to draw a line. This line is too simplistic, but it'll work, for now.
In fact I've start doing that myself. Sending PR and convincing the maintainer why the fix is necessary is just too much effort.
Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.
PopOS is not a product being sold by techbros trying to pump their stock like AI is, it's free open source software, it doesn't need to "outcompete" anything. If you don't like it, don't use it.
My repos probably don't see as much traffic as SQLAlchemy though.
Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.
So basically you're saying you reject drive-by PRs.
I have a hard time knowing if anti AI is a mental illness or propaganda coming out of China.
Unironically.
Separately, why would anyone use a Debian based desktop OS? Your $11 Amazon mouse won't work. An Nvidia card won't work. Just use Fedora.