This is a pretty good middle ground, I think. You can't prevent LLM usage and there's significant downsides to doing so universally, so restricting contributions to things that a human needs to demonstrably understand circumvents a lot of problems.
I think this should be combined with banning people who cheat by trying to explain code without properly reviewing it and burning cycles from humans at the other side.
At least that would be my policy if AI is allowed.
For until the end of the article I was thinking GCC as in Gulf Cooperation Council. No idea why, but I was surprised this being about the gcc compiler and AI rules.
>Why is anyone not surprised that some folks are not enthusiastically building their own gallows?
Their way of putting it is funny. I like this person's opinion, but I disagree with it. It's just their own framework, but I think it could also serve as a foundation for building other things.
Speaking of PRs, honestly, I've done the same thing before—it was just a one-line fix, but I asked the LLM to add 30 lines of tests just to look more professional
It’s definitely interesting watching the OSS and commercial world swing in seemingly opposite directions on this.
It would be nice to see some companies sharing more balanced successful practices they’ve implemented with AI
I don't use gcc directly - not in a long while - but almost everything I rely on uses it, and it's hugely encouraging to have the stewards of this project contemplate, and then determine to have this policy.
Meanwhile, I don't know who quotemstr is, but they don't sound sane in any of the exchanges in this thread.
How is this encouraging? They're sticking their heads in the sand and dooming themselves to irrelevance. All but the most strongly and wrongly ideologically motivated will contribute to other projects like LLVM when GCC asks them to code with rocks and sticks instead of taking advantage of arguably the most important invention in human history.
You threaten to ban any existing contributors lying about their code and you increase scrutiny of new contributors. You can go a step further and reach out to any other projects that the lying contributor works on and let their leadership know about the lying. There's plenty of ways to use social pressure that you're pretending don't exist. Not everything is a code problem. Some things are people problems and we've got about half a million years of experience with that.
I wonder how they plan to detect it something is LLM generated. I think what this leads to is people just working hard to make their outputs appear human generated.
It's extremely disturbing to see literally all foundational projects succumbing to the slop-monster. It's only a matter of time now until the critical mass of hard to detect bugs accumulate in the project, making gcc completely unusable for any practical purpose.
What's worse - those will be subtle bugs The kind you get from having a defective RAM chip, somewhere in the upper addresses.
And if we can't trust the compiler, we can't trust anything compiled with it.
It's always fun to see the people throwing fits over this kind of thing. No matter what approach/wording they use, and no matter how hard I try to give them the benefit of the doubt, the mental images my mind forms of these people is always entertaining.
Ofc, it's less fun to accept that many of them are probably bots, but whatever.
I think if the submitter can answer questions about the code, and exhibit understanding for every line then it should be indistinguishable. But I don't maintain any busy projects.
The moderating should focus on good user participation, and a reputation to give old users leeway. I'd be as specific as requesting new users to respond as succinctly as possible to avoid AI ranting
Why do the AI bros even care? Surely you can just make a better gcc with a prompt right, why care about one project disallowing your Thoughtful Contributions?
Makes sense. The G in GCC is for GNU right, GNU as in Stallman-style Free Software. The GPL operates based on copyright licenses. If LLM output can not be copyrightable (as the courts seem to assert), then it can not be a significant part of Free Software.
But looking at the history of the free software movement, it seems like they should actually be embracing LLMs. It's interesting how differently people think.
The starting point of GNU was that Unix was expensive and costly for research labs, so they set out to build a free alternative that users could control from the ground up.
So if LLMs are useful, shouldn't we be building a free LLM ecosystem where users can run, study, and modify them, rather than letting a few companies control access to models, execution, environments, and data processing?
Of course, it's natural for organizations to drift from their original mission as they get older.
But judging by GNU's early history, the logic that:
1.LLMs themselves are bad because companies control them,
2.Writing code with AI isn't real programming,
3.Only human-written code is truly free.
This logic seems a bit flawed. After all, compilers, debuggers, and automated builds all automated tasks that humans used to do manually. And the GNU project itself created tools like Make and GDB so that programmers could work at a higher level.
If LLMs can reduce repetitive coding, documentation browsing, translation, test generation, and understanding legacy code, then that seems perfectly aligned with the next goals of free software. Making knowledge accessible to more people rather than keeping it locked up as tacit knowledge held by a few experts.
I guess when organizations grow large, they inevitably attract people who don't fully align with the original purpose
The starting point of GNU was that Stallman was pissed he got in trouble when he got caught copying code from the Symbolics sources to the MIT and LMI sources, which was against the agreement Symbolics and LMI had with the AI Lab, which was that improvements could only flow one-way (AI Lab to commercial). Dan Weinreb (RIP) confirmed this publicly.
Of course, not long after starting GNU, Stallman got caught copying code from Unipress emacs sources into the then-new GNU emacs sources. Oops! That’s why it was difficult for quite a long time to find early GNU emacs sources online—they were purged from various archives because they were infringing.
Without wanting to take a stance on either side of these arguments, it does occur to me that, for this and other major FOSS projects deciding on AI policies, others can always start forks with different policies. The success or failure of such forks might even offer some insight into how helpful or harmful different forms of AI use are, at least from a programming perspective.
81 comments
[ 0.22 ms ] story [ 21.9 ms ] threadI think this should be combined with banning people who cheat by trying to explain code without properly reviewing it and burning cycles from humans at the other side.
At least that would be my policy if AI is allowed.
> We welcome all contributors to the community even if they have not yet followed our policies; we should guide such contributors on how to do so.
Kudos to the GNU project for their attitude.
Their way of putting it is funny. I like this person's opinion, but I disagree with it. It's just their own framework, but I think it could also serve as a foundation for building other things.
Speaking of PRs, honestly, I've done the same thing before—it was just a one-line fix, but I asked the LLM to add 30 lines of tests just to look more professional
Once the 3 big ones start using LLM to review/accept the work for speed reliability sake, who knows what is going to happen.
Meanwhile, I don't know who quotemstr is, but they don't sound sane in any of the exchanges in this thread.
Don't be ridiculous.
at some point, what's stopping people from lying or make the code like human writing one ????
(Extra fun if the AI generated compiler is under BSD licence.)
Ofc, it's less fun to accept that many of them are probably bots, but whatever.
The moderating should focus on good user participation, and a reputation to give old users leeway. I'd be as specific as requesting new users to respond as succinctly as possible to avoid AI ranting
GCC is much less relevant than LLVM in the AI world. NVIDIA's entire CUDA compiler stack is built on LLVM, just like AMD's HIP stack.
The starting point of GNU was that Unix was expensive and costly for research labs, so they set out to build a free alternative that users could control from the ground up.
So if LLMs are useful, shouldn't we be building a free LLM ecosystem where users can run, study, and modify them, rather than letting a few companies control access to models, execution, environments, and data processing?
Of course, it's natural for organizations to drift from their original mission as they get older.
But judging by GNU's early history, the logic that:
1.LLMs themselves are bad because companies control them,
2.Writing code with AI isn't real programming,
3.Only human-written code is truly free.
This logic seems a bit flawed. After all, compilers, debuggers, and automated builds all automated tasks that humans used to do manually. And the GNU project itself created tools like Make and GDB so that programmers could work at a higher level.
If LLMs can reduce repetitive coding, documentation browsing, translation, test generation, and understanding legacy code, then that seems perfectly aligned with the next goals of free software. Making knowledge accessible to more people rather than keeping it locked up as tacit knowledge held by a few experts.
I guess when organizations grow large, they inevitably attract people who don't fully align with the original purpose
Of course, not long after starting GNU, Stallman got caught copying code from Unipress emacs sources into the then-new GNU emacs sources. Oops! That’s why it was difficult for quite a long time to find early GNU emacs sources online—they were purged from various archives because they were infringing.