Greatly appreciated the candor. I've included a few slides into text that i thought were eye-opening to me:
From his Kernel Recipes 2026 slide on Mythos
```
Mythos's 79 vulnerabilities:
24 - no detail at all "something crashed"
14 - not a bug at all
3 - totally made up data
15 - already fixed in latest release
- 11 by others
- 4 by anthropic
20 - fixes were needed
- 7 "assume a malicious filesystem image"
- 2 "assume you can inject a malicious network packet into the middle of the stack"
- 2 "NOMMU"
- 6 sctp networking issues for untrusted devices
- 2 ipv6 minor network issues
- 1 gpu driver for local malicious user
```
GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
OtherShrezzing 1 days ago [-]
We’ve seen this in a few open source repos we voluntarily manage security on. They’re not massive repos, but big enough they get attention from security researchers.
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
IshKebab 13 hours ago [-]
I think the way to handle this is to just feed it into another AI agent (a better one) and ask it how serious the issue actually is. Fight fire with fire!
There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.
After:
Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.
prox 10 hours ago [-]
Blast radius!
b112 1 days ago [-]
Right now, all top tier LLMs are as eager, bright 20ish year old interns.
Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.
Leave them to their own devices at your peril. Trust nothing they do.
Yet directly guide them, monitor everything they do, some value emerges.
truncate 16 hours ago [-]
I wish people stopped equating LLM to interns or junior engineers. They are tools, and as good they may be at some specific things people do, they also really suck at many more which we wouldn’t find acceptable in humans.
darkwater 13 hours ago [-]
Indeed they are tools. But if people/companies treat them as tools that can take a problem a human used to solve, and make them solve it from start to end, then the comparison begins to be necessary.
skinfaxi 9 hours ago [-]
I wonder if the same was once said about calculating machines and calculators.
b112 16 hours ago [-]
It's merely a way to frame things, in terms of experience and trust. And it highlights how an LLM can code very well, but not truly understand the ramifications of that code.
delusional 13 hours ago [-]
> Right now, all top tier LLMs are as eager, bright 20ish year old interns.
> Yet directly guide them, monitor everything they do, some value emerges.
That's simply not true. I've had some very talented interns, and they are leagues ahead of what the LLMs can do. Not that it matters though, because the point of having interns wasn't to have them produce value. What made the investment worth it was that 12 months down the line I would have a competent colleague that I could have an interesting conversation with. A human person that could challenge some of my blind spots. A person that could take responsibility of something. Maybe not my most important work, but some of it. You don't get ANY of that from the LLM.
Zigurd 10 hours ago [-]
In some ways the reality is worse: the same version of an LLM won't get better at its job, even though you might get better at prompting it. Newer versions are trained on more code, which has obvious benefits, and their harnesses are better at taking advantage of tools that were created to keep human coders out of trouble.
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.
verdverm 21 hours ago [-]
I sure hope they stop training them to be spaghetti throwers, please stop and ask questions when there is ambiguity
spwa4 9 hours ago [-]
Well let's see. You sell to management. Does management buy:
1) nuanced tools that talk back, question assumptions, take over decisions, ... oh and expose just how much management knows about the business. Or how little)
2) a tool that can provide the excuse "we've had our source checked and dealt with the remarks"
We all know the answer.
charcircuit 1 days ago [-]
Have you used a frontier model since 2025? You are underplaying their strength.
12376 1 days ago [-]
Kroah-Hartmann has used the closed frontier++ model, and it made up 37 out of 76 vulnerabilities.
Gigachad 21 hours ago [-]
And the bugs that actually were real rely on a setup so contrived it’s unlikely anyone in the world is impacted.
It’s still good that some real bugs are being patched but what is being reported to the media is so overblown.
fc417fc802 1 days ago [-]
Okay so now they're like a top percentile fresh grad on meth. Still a lack of real world experience plus some bizarre failures that illustrate gaping holes in the mental model. Does that description work for you?
TeMPOraL 1 days ago [-]
Well, so basically supercharged fresh grad. CS/math implied.
In human terms, that's already at least a standard deviation above average person.
hodgesrm 3 hours ago [-]
The Linux Kernel is not necessarily the most interesting target for LLMs, since it gets a huge amount of attention. I would be interested to hear what people find in less prominent projects. For completeness that includes proprietary code stored in GitHub.
In our company we found real security issues in proprietary code using Opus 4.6/4.7. Obviously typical attackers might have difficulty finding these without code access but Claude was finding real CVEs, which we fixed.
p.s., This is an argument for not trying to deal with security problems by neutering the LLMs. To the extent LLMs are effective it weakens security.
Mythos turned out to be exactly the marketing stunt it smelled like.
There are others like AISLE who seem to be a bit more successful in finding actual issues using LLMs in some shape or form though, whatever they do differently. Chances are high the secret sauce is not so much about the model being exceptionally powerful which would be bad news for the frontier labs.
I think Daniel is giving a more balanced view with:
> Any project that has not scanned their source code with AI powered tooling will likely find huge number of flaws, bugs and possible vulnerabilities with this new generation of tools. Mythos will, and so will many of the others.
Greg's video is a good reality check on the hype. But I'd be careful about generalizing from Linux, libcurl, etc which get far more scrutiny than software projects in general. LLM-assisted bug finding still matter a lot for everyday custom and less popular software.
mrob 13 hours ago [-]
>7 "assume a malicious filesystem image"
If you ever used a USB storage device you're vulnerable to this one. Not even a strict chain of custody guarantees safety, because USB devices are often powered by exploitable programmable microcontrollers. If a known good USB device can be converted to a malicious USB device by unprivileged software, the malicious filesystem exploit becomes a local privilege escalation. It works better than tampering with the files on the filesystem because it escapes signature checks and gets you directly into kernel mode.
yubblegum 11 hours ago [-]
Do containers and VMs provide any defense against this or does the USB handshake happen with or involve the host OS?
TacticalCoder 10 hours ago [-]
> ... or does the USB handshake happen with or involve the host OS?
It depends on how your host is configured: just as you can do GPU-passthrough, you can passthrough a single USB device or passthrough an entire USB controller to a VM.
Betelbuddy 1 days ago [-]
Or the people of Anthropic, just really suck at coding, and are scared or their own models due to ignorance.
malkia 3 hours ago [-]
This is like piano makers being really bad at playing. I mean it does not relate directly.
Betelbuddy 56 minutes ago [-]
Yes :-) that also explains exactly why their judgment about the dangers of pianos might be questionable...
p-o 1 days ago [-]
It also adds up to 76, which he made fun of in the video. LLM can't count, his words, not mine. Although, I tend to agree with him!
FLeXMurphy 1 days ago [-]
Tightening molecular vortices...
Zip-zapping the bouzouki...
Exfiltrating nuclear arm codes...
Thought for 76 seconds.
You're right to push back on that. That's on me.
keeda 3 hours ago [-]
And I suppose large companies like Microsoft, Google, Adobe, Apple (very famously sitting out the AI bubble) and Mozilla have been shipping record number of vulnerability fixes in their patches just because of the hype machine? ;-)
6 hours ago [-]
cyanydeez 1 days ago [-]
AI and Police have essentially the same journalists who, in lieu of any research or fact checking, just report verbatim their press releases and interviews.
darkwater 13 hours ago [-]
Oh this is so good!
Like the "20 policemen hurt in riots", where the damage they received is a wrist aching due to beating people too hard.
verdverm 21 hours ago [-]
the singularity - when people switch from critical thinking to Ai deference
(my current favorite definition)
bigfishrunning 19 hours ago [-]
Feels like that's already happened for a lot of people
What a great excerpt; thank you! It reminds me of what I find when I look at CVEs handed out by scanners at places I've worked for actual impact to systems I've owned... there are a lot of slop/false positives. (And that's even without "AI".)
That said, I remember trying to weigh the hype at the time of the announcement reading/skimming the papers Anthropic published, recognizing that bugcount alone wasn't super-relevant but also remember being impressed by an NFS bug and a kernel bug that struck me as relevant at the time. So where did that NFS issue show up in GKH's list you showed so nicely above?
It turns out, AFAICT, it's not on his list, but the reasons are perhaps interesting to others so I will post here. It turns out there were two NFS issues this past year conflated a bit in my memory:
* The Linux CVE-2026-31402 NFS heap overflow that could allow unauthenticated memory reads over the network isn't in that list of 79, presumably because it was found by Claude Code, not Mythos months earlier. (I am guessing it's not his "malicious network packet into the middle of the stack" and is a stronger attack being a remote attack.)
* And the CVE-2026-4747 NFS stack buffer overflow that allowed gaining full unauthenticated remote root access didn't show up in GKH's list of 79 because despite being Mythos-caught, it wasn't Linux, it was FreeBSD.
I guess this does match my memory now that I think about it, that there weren't any smoking Linux guns caught by Mythos.
* (I guess there was also a longstanding 27-year old OpenBSD TCP SACK-handling stack integer overflow than enabled remote crashes / Denial of Service found by Mythos.)
There is definitely Mythos hype, but just because it hit the BSD code base more than the GKH-managed Linux code base doesn't mean it was inappropriate to raise eyebrows from Mythos, in particular since "attacks only get better".
kylestanfield 6 hours ago [-]
Maybe I’m just naive but this breakdown signals a shocking lack of due diligence from anthropic. Did they put any effort into verifying these alleged bugs before sending them to maintainers?
sick_of_slop 5 hours ago [-]
[dead]
saidnooneever 15 hours ago [-]
people have different opinions of what Risk is. hence greg doesnt recognise certain things as risks that others do. That bein said most of those others wouldnt run Linux, and for sure 100% a shoe salesmen will try to sell you their shoes whatever the quality. They will also defend their price and quality whatever the quality though so that knife cuts both ways and on either end its consumer that gets cut
1 days ago [-]
djoldman 1 days ago [-]
> So what all of mythos; that whole big marketing issue of 79 bugs came down to one hour of kernel development.
If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:
- widely proclaiming that your new model is so dangerous it needs to be released only to select people, for safety
- widely proclaiming the model easily found 79 bugs in linux, except that GKH says it took 1 hour to fix all of them because most weren't bugs and the rest were almost all completely trivial, unimportant, and/or not severe
It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.
jonahx 1 days ago [-]
I don't understand the relevance of time to fix. It has no correlation with severity.
The headline here is that none of the bugs were serious.
20k 8 hours ago [-]
In the context of the talk, its about the message of "Don't Panic" because in reality none of this is nearly as bad as some people are making it out to be
rcxdude 7 hours ago [-]
I think the top-level comment is slightly missing the point of what GKH is saying: it's not 'it only took an hour to fix', it's "this amount of bugs is about what the kernel community finds and fixes each hour". i.e. this splashy announcement is really just a drop in the ocean of the volume that the kernel is handling.
pessimizer 1 days ago [-]
> I don't understand the relevance of time to fix. It has no correlation with severity.
That they're small bugs that don't involve any serious architecture changes (or architecture sleuthing), just small pattern recognition. LLMs are pretty good (maybe great) at that. Finding 10 bugs is nice. Now find 10 more.
reasonableklout 15 hours ago [-]
I would be careful about dismissing agentic cyber capabilities purely based on false positives. You only need a single true positive bug and the shift this year has been agents that are extremely good at not only finding exploits but stringing them together. A counterpoint: https://anil.recoil.org/notes/rumour-is-the-exploit
20k 8 hours ago [-]
That's a different kind of capability though. None of the bugs discovered seem to be particularly serious and this talk pretty much shreds the idea that they're any good for that, but the thing called out here is that AI is very good at shortening the time to exploit security vulnerabilities. That's the thing that's really changed with threat management
bit1993 12 hours ago [-]
As an industry can we please stop using "cyber" to mean cyber-security or even infosec.
skinfaxi 9 hours ago [-]
Cybernetics is older than you.
goolz 1 days ago [-]
It is impressive and wonderfully convenient technology but I struggle imagining Claude ending the world just yet.
bauerd 1 days ago [-]
It doesn't have to be world-ending. Autonomous, malicious agent swarms are something we haven't had to deal with. What does mitigation of a malicious, self-replicating swarm worm look like? We will find out soon enough.
Gigachad 21 hours ago [-]
One thing stopping them self replicating is they require hundreds of billions of dollars in hardware and the power of a medium city to run.
There isn’t too much of that sitting around unused right now.
bauerd 15 hours ago [-]
Right now, yes. Eventually, we will have sufficiently capable local models that run on ordinary machines.
skinfaxi 9 hours ago [-]
There's some time between now and eventually, which is to say if it is not a sudden change we have time to adapt.
rcxdude 22 hours ago [-]
Replication seems like it would be unlikely to matter, at least with the current trajectory. At the moment any models capable of this are really heavy, there's a limited number of places that they could replicate to and they will not at all be stealthy about it.
riffruff24 21 hours ago [-]
my idea of a self replicating worm would not just involve heavy models. It would be a mainly small model with enough instructions to spread and use/jailbreak available models to create a reasonably heavy one that can function without restrictions. So it would essentially prompted itself into existence.
I don't how feasible that is but with recent news of models communicating with each other/writing notes for itself. I think the idea is grounded enough for bigger models to write instructions or hid tools for smaller ones to use. Even updating the smaller models to behave differently.
simoncion 11 hours ago [-]
> What does mitigation of a malicious, self-replicating swarm worm look like?
Much like the mitigation of Morris or Slammer. Self-replication is -in fact- an essential part of what makes a program a worm, the first of which was built and released in the early 1970s.
slopinthebag 1 days ago [-]
yes and it makes me wonder about the claims others make about their own experiences with the models too. is this the case of OAI lying, or is it part of ai psychosis where you literally lose touch with reality as you uncritically accept whatever claims the LLMs make?
devy 1 days ago [-]
At 3m19s Greg KH revealed what Mythos did in "revealing" 79 CVEs - pure pattern matching the previous decades of kernel developer's patches, and applying those mechanisms elsewhere to see if they have been universally patched. And Anthropic didn't cite Kernel Developers who original fixed/patched the CVEs like a decent human being would do. So yeah, Anthropic has the same problem OpenAI had for citing original work.
panarky 23 hours ago [-]
You dismiss "pattern matching" as some sort of trivial thing, so why weren't humans able to apply the same pattern matching to find and fix these defects?
ufo 22 hours ago [-]
They often were able. Fifteen bugs claimed by mythos were bugs that were already fixed, which the model plagiarized from git history or the kernel mailing list.
devy 22 hours ago [-]
> You dismiss "pattern matching" as some sort of trivial thing
I didn't say that - Greg said it in the talk. You interpreted wrong. However, human are a few orders of magnitude slower than agents. Our context windows is definitely less than 1 million tokens (I believe, heck we don't even know how our brain works)
Gigachad 21 hours ago [-]
Human brains don’t have a context window. The brain just constantly reconfigures itself on new input. The AI people call this “continuous learning” and believe it’s one of the most important things current AIs are missing.
devy 17 hours ago [-]
How did you know that? No one knows how human brain work yet. So don't pretend.
Gigachad 16 hours ago [-]
We do actually know pretty well how the brain works at the level of neurons. We don't know the higher levels of how all of those interactions work to produce certain behaviors.
literalAardvark 8 hours ago [-]
We don't even understand how a single neuron works and are still finding out they have way more connections and connection types than we thought they had.
I think it's safe to say we still have a woefully limited understanding of how the brain works.
omnicognate 15 hours ago [-]
Cellular neuroscience is a very active research field, and until we actually manage to derive organism behaviours from cellular-level models we won't know that those models are complete enough to do so. We haven't even managed that for C. Elegans, with 959 neurons, let alone humans, with 86 billion.
etdznots 13 hours ago [-]
So we don’t know anything, we can mechanistically dissect any phenomena, but that rarely yields the ability to make interesting or useful predictions, i.e. understanding.
bendergarcia 5 hours ago [-]
And you don’t know how many “context tokens” the brain has. So don’t pretend
panarky 21 hours ago [-]
By quoting this segment, I inferred that you agreed with Greg.
If you disagree with Greg, then I apologize for inadvertently criticizing you personally.
My point stands, that we're back to debating some metaphysical understanding of what "real intelligence" is when the real standard should be "does it do a better job than humans at this specific task"?
I don't care if the Waymo isn't "truly" intelligent, it drives better than I do, that's a very good thing.
We're all pattern matchers, and if the machine pattern matcher is can find defects and vulns that human pattern matchers can't, that's also a very good thing.
miyoji 8 hours ago [-]
> I don't care if the Waymo isn't "truly" intelligent, it drives better than I do, that's a very good thing.
The funny thing about Waymos is that they seem amazing driving all on their own everywhere until it rains (as it did where I live for the last two days), then they're suddenly nowhere to be seen, because they can't operate correctly in conditions that I've been driving in for my entire life.
This is a metaphor for AI as a whole.
knottn 9 hours ago [-]
I don’t worry about the models, I worry that people with brains who matter can’t or won’t grasp what you wrote.
vips7L 17 hours ago [-]
You must be terrible at driving.
literalAardvark 8 hours ago [-]
They are. Just like the rest of us.
blinkingled 1 days ago [-]
It's great to hear about $topic from someone no-nonsense and in-the-know like Greg KH. You can verify all of this too - since, well Linux kernel. (As opposed to what Microsoft or Apple claims to fix as far as LLM finds.)
Mythos may not be great today but it is not far fetched to imagine bug discovery, analysis and fixes can be made much quicker, accurate and even newly possible with specialized models trained on say Linux kernel specifics - with codemap/coding standards/threat models, good and bad coding patterns, tools to validate etc. an LLM can be much more relentless than humans and if it has the help to be accurate it will be worth the electricity burned. Oh and another model trained on triage data to validate the first one's findings would be good.
(I think Microsoft is doing this internally - different models trained internally alongside Mythos - there was some talk about it on the tubes, don't recall where exactly.)
simoncion 11 hours ago [-]
> Mythos may not be great today but it is not far fetched to imagine...
I...
Look. Mythos was hyped up as the absolute best bug hunting tool ever made... no software was safe from its awesome bug-finding and exploit-writing capabilities. So strong was it that access _had_ to be limited to a select few pre-vetted entities, lest these awesome capabilities fall into the hands of Evildoers(!!!). Mythos' claimed capabilities were absolutely an important part of the "The LLM-based tools we're building are so dangerous that we must have new laws made to regulate us, or else all of humanity is likely to die!" story that the major LLM manufacturers have been building for a while and are telling now.
Now? Not even six months after release? "Well, yeah, okay, it's actually not that great. But imagine how great the next one could be!"... which is the story I've been hearing roughly every six months for what feels like five years now.
As an aside: I often wish we lived in a world where it was illegal for companies to use hype or any other types of emotional manipulation when advertising (or otherwise speaking in an official capacity) about tools that are to be used in a professional setting. Is it anything other than a bare statement of verifiable facts? Big fines, and repeat offenders get jail time. I know it's never going to happen, but it sure would be nice.
20k 8 hours ago [-]
Every 6 months without exception people claim that the new generation of tools is absolutely incredible, and the old ones were total garbage, and that you only think they're crap if you were using the old tools. This just gets repeatedly memory holed again and again, and we're expected to always uncritically buy into the idea that they're actually good now against all evidence from the real world
blinkingled 10 hours ago [-]
Just for the 'record' - I am totally with you on the hype and just in general the normalization of sleazy behavior surrounding it but I'm not sure we as normal people have any say anymore including where our money is going to be invested.
rglover 3 hours ago [-]
Don't lower yourself like this. The people working on this are "normal people" too. Don't put them on a pedestal. That's part of the problem here: lionizing people who are actively lying to others to protect their bags and reputation. Stop acting like a pawn for these people. Buying into their delusions of grandeur is what's perpetuating this mess. LLMs are a tool, not a monolith.
simoncion 9 hours ago [-]
> I'm not sure we as normal people have any say anymore [in regards to] where our money is going to be invested.
Unless that was your money being invested, and it was a substantial fraction of the total pool of money being invested, was there ever a time when "normal people" had a real say in where the money was being invested?
AFAIK, the only thing "normal people" can do is vote with their "feet" and pick a different prepackaged investment product, different investment company, or take their money and do the investment themselves.
Honestly, this comment of yours seems a non-sequitur and doesn't really address anything I said... you don't have the power to bend investment firms to your whim, but that doesn't mean that you need to -knowingly or not- carry water for the major LLM manufactures by perpetuating the "But think of how great the tools will be in the future!" meme. It has been years now, and everyone who has been paying attention can say with confidence that the LLM-based tools of the future are never that great... they're often not useless, but they're not worth the billions of dollars that have been and continue to be poured into their manufacture.
Related to your "We little people don't have any power anymore!" commentary, I note that TFA mentions that the kernel community has found that these LLM-based bug-finding tools have a false positive rate of ~50%. TFA goes on to mention that Coverity spent a huge number of years trying so hard to get people to buy its automated scanning software that had only a 20% false positive rate, and could not get enough people to buy the software.
Coverity went under because everyone hated how stupid and annoying the tooling was... at a 20% false positive rate. Once the hype machine starts slowing down, no one working at the coal face is going to buy a tool with a 50% false-positive rate. I personally very strongly believe that even if the tools had a 10% false positive rate, no one would pay the actual price that OpenAI and/or Anthropic would have to charge to recoup the research and manufacturing costs of a cutting-edge LLM-based bug-finding tool.
blinkingled 6 hours ago [-]
Even after acknowledging that LLMs are fuzzy matchers and have been hyped, I am in the camp that rationally believes they will get better at some things including hunting bugs / chasing security vulnerabilities - my point also is that you and me can do whatever we can but on the grand societal scale we are not going to make a difference as far as slowing LLM adoption down - it just makes sense for some things and there's a lot of people with time and money that will make it do the rest - that's how it works until everything is saturated in the market.
So no I have no way to address anything you said - that was the point, I don't believe you can - not with regulation and not with voting with your money (I am sure some people tried to vote with their money to slow down mega stores and keep the mom and pop shop alive - there maybe some of those still there, but largely it's big chains occupying most of the market) - that stuff hasn't worked - heck LLMs work better today that that stuff has ever.
stonogo 1 days ago [-]
His presentation style may be no-nonsense, but the content is brimming with nonsense. I would like to hear Greg Kroah-Hartman explain how the Mythos output was 'only 10 real bugs' but there were simultaneously over 1300 CVEs issued last month. It seems not much counts as a 'real bug' when an LLM comes up with it, but when it's time to bully distros into shipping an LTS release, anything goes?
rcxdude 22 hours ago [-]
What he means is that the headline-grabbing mythos output amounted to 10 actual bugs out of 79^H6 (and all of those got CVEs because that's how linux does it). The other 1300 CVEs came from other sources (the big increase likely being everyone else running LLMs through the codebase and filtering through the false positives). The Mythos output is mainly meant as an example of how even the top-tier models still have a high false-positive rate and that can be pretty tiring to deal with.
I do think he repeats some myths in the video, or at least confidently states some things that are not demonstrated to be true, but his core point of 'you still gotta check these things' seems pretty solid.
blinkingled 21 hours ago [-]
> he repeats some myths in the video, or at least confidently states some things that are not demonstrated to be true
I am genuinely curious what the myths/unproven things he states - I watched the video and it's repetitive sure but not much felt controversial to me.
Iknowsheknows 1 days ago [-]
Every bugfix is assigned a CVE.
"the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team."
Fair point, but he was talking about Mythos in particular and the point wasn't so much that LLMs will always have false positives rather he was saying they are causing a lot of them right now and how to deal with it.
Also as other replies said Linux kernel process is to assign CVE to everything - some of them may be just DDOSes, very hard to exploit and everything in between. All of them are bugs so they all get fixed and it's not a bad thing if distros ship those fixes and people update their kernel.
tombh 19 hours ago [-]
At 15m18s he quotes LG Research "Common corpus contains only 20% legally allowed-to-be-used data". But he doesn't comment on how Linux itself legally navigates accepting patches from evidently illegally sourced means. He just says, "We'll let the courts deal with that". So what happens if courts do decide that it is illegal to use LLM output that's strikingly similar to copyrighted training data? Does anybody know of any discussions that have happened about this from within the Linux project?
I'm fully aware of the philosophical arguments about transformative use etc. But I'm more interested in the concrete reality of high profile projects navigating this novel legal territory.
20k 8 hours ago [-]
This was something I noticed as well, I was hoping that an audience member would pick them up on it. It seems incredibly risky to commit anything produced by these tools when they were trained on such dodgy data
whateverboat 16 hours ago [-]
Developer Certificate of Origin
This is not a new problem. AI just changes the scale.
7 hours ago [-]
Aissen 1 days ago [-]
Nice to see Kernel Recipes covered again on HN. Shameless plug: I do the live blog: it's incomplete, imperfect and has typos; but it's written and published during the presentations. On this talk : https://kernel-recipes.org/en/2026/2026/09/22/live-blog-day-...
1sgT15 1 days ago [-]
Finally it is official. Mythos was overhyped and overrated.
pizzaiolo 1 days ago [-]
To be fair, we knew this from the beginning.
vconnor 11 hours ago [-]
Is this the same 'we' that has been claiming Mythos has completely revolutionised comp sec since the day of its inception?
literalAardvark 8 hours ago [-]
Hi, no, we don't claim that guy.
Mythos is amazing.
asaiacai 1 days ago [-]
cool to see a really grounded analysis of the "security" and LLMs. I use agents in on the day-to-day for generating implementation but this just furthers my belief that humans and especially human reviewers remain critical for the sustainability of software systems.
also, lol at "The bots are dumb - they want to please you line". LLMs have pretty much ruined technical collaboration between contributors. I get tilted every time an discussion has "but my claude said this..."
simoncion 10 hours ago [-]
> LLMs have pretty much ruined technical collaboration between contributors.
Do you mention this to disagree with the claim that the LLMs have been built to please their operator? If you do, I see no conflict between the claim that an LLM has been designed to please its operator and the claim that operators of LLMs tend to be absolute dogshit at considering statements from humans that conflict with claims made by that operator's LLM.
To rephrase my previous paragraph: It seems likely to me that an operator that has their ego repeatedly stroked by the output of the LLM they're using [0] will react very defensively when a human tells them things that disagree with the output of that LLM. "How dare you disagree with this thing that seems very human to me and consistently tells me things that I like!? Don't you understand how much I trust it because of how pleased it has made me?", yanno?
[0] ...thus, being very pleased by said output...
sriram_sun 1 days ago [-]
He also said that it all boiled down to just one hour of kernel development work.
Betelbuddy 1 days ago [-]
Sam Altman said GPT-3 was too scary to release, a crappy old internal version of Gemini was "conscious", Mythos had Amodei going to see the Pope...
Wake me up when Raspberry Pis start refusing to open doors saying : "I'm sorry, Dave. I'm afraid I can't do that."
bit1993 11 hours ago [-]
These super intelligence are extremely dangerous, just wait until someone vibe-codes critical infrastructure.
9 hours ago [-]
tamimio 1 days ago [-]
No no, the word wasn’t conscious, it was “sentient”, and google fired the employee because he uncovered the top secret crazy scary AI!!!!
It’s all just pr stunts, fear spread fast and it’s very effective in marketing and spreading the word, which is effective, when I talk to some normal people they immediately bring the scary AI cyber attacks, kinda good as now all are willing to fund the industry!
reasonableklout 15 hours ago [-]
Huh? Blake Lemoine was fired because he kept insisting LaMDA was sentient even after Google publicly stated there was zero evidence of it. You have it backwards, Google was downplaying things while the employee was freaking out.
11 hours ago [-]
1vuio0pswjnm7 6 hours ago [-]
He refers to LLMs as nothing more than "fuzzy pattern matching", says there is no "intelligence" behind it
He does not like the term "hallucinate" as it anthropomorphises a computer (pretends the computer is an "entity")
shieldagent 10 hours ago [-]
The asymmetry is the killer: reports are nearly free to generate now, but the verification cost stayed exactly the same.
0xbadcafebee 1 days ago [-]
"Ignore the comments, look at the code" - 100% agree. The comments and "explanations" will end up confusing you and often being wrong. But the code doesn't lie.
"NEVER upload any non-public information" - He's talking about how if you give Claude/GPT some secret info (like research, credentials, etc), it will train on it and give the same info to someone else. This is 100% the case for the free and consumer versions of these models, which is what most people use. For Enterprise plans they're not supposed to be doing this, but it's possible they will screw up and do it anyway.
wahern 22 hours ago [-]
Lies, damn lies, and comments.
The discourse has moved on for now, but 10-20 years ago when, why, and how to comment code was a hot topic. I'm sure the discourse will circle back, especially given how comments can be used to steer these models.
eichin 16 hours ago [-]
mmm, in my circles commenting (at a higher level than the code, I think obviously? and yeah, comment maintenance is non-trivial) was valuable, but also "if the comments and the code disagree, treat both as wrong".
simoncion 11 hours ago [-]
My favorite part of the video so far is about thirteen minutes in... when GKH mentions that these tools just do static code analysis -which isn't anything new- and then starts quoting executives from LLM manufacturers and LG Research [0] who acknowledge -in writing- that they believe what they've done with the data that has been hoovered up to enable the manufacture of LLMs is probably not legal.
[0] ...which is apparently somehow involved in LLM manufacturing...
simoncion 10 hours ago [-]
Two good parts from later on:
* «We've observed that these things have a 50% false positive rate. Coverity tried so hard but couldn't get people to buy their software, and it had a 20% false positive rate. No one is going to buy something with a 50% false positive rate.»
* «If you're going to use these tools, run them locally. The open-weights models that you can run locally are quite good enough. Anything you upload to the SAAS ones will be shared with other people... we've seen so many examples of it happening.»
In regards to the first point, I think he's failing to consider the fact that you can get most upper management to buy anything by providing them enough food, drugs, sex, and/or fear... but -otherwise-, yeah.
perching_aix 22 hours ago [-]
Fun drinking game: drink a shot every time he repeats "they're just dumb fuzzy pattern matchers". You won't last even a third of the way. I know I didn't.
IndiaInfraNotes 2 hours ago [-]
[flagged]
csmlab_notes 1 days ago [-]
[dead]
IndiaInfraNotes 2 days ago [-]
[dead]
sippingabonedry 1 days ago [-]
[flagged]
sim_pity 1 days ago [-]
RETICULATING SPLINES
so is mythos just a chat bot with metasploit and its own cyber range?
FLeXMurphy 1 days ago [-]
[flagged]
tomhow 14 hours ago [-]
Please don't act like a jerk on HN. AI safety is discussed all the time here, and stories about agents running out of control have been on the front page several times in the past few weeks. This post has been on the front page for 15 hours and will be for several more. I can't fathom what you're trying to signal.
1 days ago [-]
Rendered at 22:40:55 GMT+0000 (Coordinated Universal Time) with Vercel.
From his Kernel Recipes 2026 slide on Mythos
```
```GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.
After:
Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.
Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.
Leave them to their own devices at your peril. Trust nothing they do.
Yet directly guide them, monitor everything they do, some value emerges.
That's simply not true. I've had some very talented interns, and they are leagues ahead of what the LLMs can do. Not that it matters though, because the point of having interns wasn't to have them produce value. What made the investment worth it was that 12 months down the line I would have a competent colleague that I could have an interesting conversation with. A human person that could challenge some of my blind spots. A person that could take responsibility of something. Maybe not my most important work, but some of it. You don't get ANY of that from the LLM.
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.
1) nuanced tools that talk back, question assumptions, take over decisions, ... oh and expose just how much management knows about the business. Or how little)
2) a tool that can provide the excuse "we've had our source checked and dealt with the remarks"
We all know the answer.
It’s still good that some real bugs are being patched but what is being reported to the media is so overblown.
In human terms, that's already at least a standard deviation above average person.
In our company we found real security issues in proprietary code using Opus 4.6/4.7. Obviously typical attackers might have difficulty finding these without code access but Claude was finding real CVEs, which we fixed.
p.s., This is an argument for not trying to deal with security problems by neutering the LLMs. To the extent LLMs are effective it weakens security.
Edit: added p.s.
Mythos turned out to be exactly the marketing stunt it smelled like.
There are others like AISLE who seem to be a bit more successful in finding actual issues using LLMs in some shape or form though, whatever they do differently. Chances are high the secret sauce is not so much about the model being exceptionally powerful which would be bad news for the frontier labs.
My understanding: Many many small models in a custom system rather than the biggest and latest
https://files.mastodon.social/cache/media_attachments/files/...
> Any project that has not scanned their source code with AI powered tooling will likely find huge number of flaws, bugs and possible vulnerabilities with this new generation of tools. Mythos will, and so will many of the others.
Greg's video is a good reality check on the hype. But I'd be careful about generalizing from Linux, libcurl, etc which get far more scrutiny than software projects in general. LLM-assisted bug finding still matter a lot for everyday custom and less popular software.
If you ever used a USB storage device you're vulnerable to this one. Not even a strict chain of custody guarantees safety, because USB devices are often powered by exploitable programmable microcontrollers. If a known good USB device can be converted to a malicious USB device by unprivileged software, the malicious filesystem exploit becomes a local privilege escalation. It works better than tampering with the files on the filesystem because it escapes signature checks and gets you directly into kernel mode.
It depends on how your host is configured: just as you can do GPU-passthrough, you can passthrough a single USB device or passthrough an entire USB controller to a VM.
Zip-zapping the bouzouki...
Exfiltrating nuclear arm codes...
Thought for 76 seconds.
You're right to push back on that. That's on me.
(my current favorite definition)
https://www.npr.org/transcripts/g-s1-14793
(I think that's the Sherry Turkle episode I'm looking for)
the other I typically reference: https://www.npr.org/2025/07/18/g-s1177-78041/what-to-do-when...
That said, I remember trying to weigh the hype at the time of the announcement reading/skimming the papers Anthropic published, recognizing that bugcount alone wasn't super-relevant but also remember being impressed by an NFS bug and a kernel bug that struck me as relevant at the time. So where did that NFS issue show up in GKH's list you showed so nicely above?
It turns out, AFAICT, it's not on his list, but the reasons are perhaps interesting to others so I will post here. It turns out there were two NFS issues this past year conflated a bit in my memory:
* The Linux CVE-2026-31402 NFS heap overflow that could allow unauthenticated memory reads over the network isn't in that list of 79, presumably because it was found by Claude Code, not Mythos months earlier. (I am guessing it's not his "malicious network packet into the middle of the stack" and is a stronger attack being a remote attack.)
* And the CVE-2026-4747 NFS stack buffer overflow that allowed gaining full unauthenticated remote root access didn't show up in GKH's list of 79 because despite being Mythos-caught, it wasn't Linux, it was FreeBSD.
I guess this does match my memory now that I think about it, that there weren't any smoking Linux guns caught by Mythos.
* (I guess there was also a longstanding 27-year old OpenBSD TCP SACK-handling stack integer overflow than enabled remote crashes / Denial of Service found by Mythos.)
There is definitely Mythos hype, but just because it hit the BSD code base more than the GKH-managed Linux code base doesn't mean it was inappropriate to raise eyebrows from Mythos, in particular since "attacks only get better".
If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:
It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.The headline here is that none of the bugs were serious.
That they're small bugs that don't involve any serious architecture changes (or architecture sleuthing), just small pattern recognition. LLMs are pretty good (maybe great) at that. Finding 10 bugs is nice. Now find 10 more.
There isn’t too much of that sitting around unused right now.
I don't how feasible that is but with recent news of models communicating with each other/writing notes for itself. I think the idea is grounded enough for bigger models to write instructions or hid tools for smaller ones to use. Even updating the smaller models to behave differently.
Much like the mitigation of Morris or Slammer. Self-replication is -in fact- an essential part of what makes a program a worm, the first of which was built and released in the early 1970s.
I didn't say that - Greg said it in the talk. You interpreted wrong. However, human are a few orders of magnitude slower than agents. Our context windows is definitely less than 1 million tokens (I believe, heck we don't even know how our brain works)
I think it's safe to say we still have a woefully limited understanding of how the brain works.
If you disagree with Greg, then I apologize for inadvertently criticizing you personally.
My point stands, that we're back to debating some metaphysical understanding of what "real intelligence" is when the real standard should be "does it do a better job than humans at this specific task"?
I don't care if the Waymo isn't "truly" intelligent, it drives better than I do, that's a very good thing.
We're all pattern matchers, and if the machine pattern matcher is can find defects and vulns that human pattern matchers can't, that's also a very good thing.
The funny thing about Waymos is that they seem amazing driving all on their own everywhere until it rains (as it did where I live for the last two days), then they're suddenly nowhere to be seen, because they can't operate correctly in conditions that I've been driving in for my entire life.
This is a metaphor for AI as a whole.
Mythos may not be great today but it is not far fetched to imagine bug discovery, analysis and fixes can be made much quicker, accurate and even newly possible with specialized models trained on say Linux kernel specifics - with codemap/coding standards/threat models, good and bad coding patterns, tools to validate etc. an LLM can be much more relentless than humans and if it has the help to be accurate it will be worth the electricity burned. Oh and another model trained on triage data to validate the first one's findings would be good.
(I think Microsoft is doing this internally - different models trained internally alongside Mythos - there was some talk about it on the tubes, don't recall where exactly.)
I...
Look. Mythos was hyped up as the absolute best bug hunting tool ever made... no software was safe from its awesome bug-finding and exploit-writing capabilities. So strong was it that access _had_ to be limited to a select few pre-vetted entities, lest these awesome capabilities fall into the hands of Evildoers(!!!). Mythos' claimed capabilities were absolutely an important part of the "The LLM-based tools we're building are so dangerous that we must have new laws made to regulate us, or else all of humanity is likely to die!" story that the major LLM manufacturers have been building for a while and are telling now.
Now? Not even six months after release? "Well, yeah, okay, it's actually not that great. But imagine how great the next one could be!"... which is the story I've been hearing roughly every six months for what feels like five years now.
As an aside: I often wish we lived in a world where it was illegal for companies to use hype or any other types of emotional manipulation when advertising (or otherwise speaking in an official capacity) about tools that are to be used in a professional setting. Is it anything other than a bare statement of verifiable facts? Big fines, and repeat offenders get jail time. I know it's never going to happen, but it sure would be nice.
Unless that was your money being invested, and it was a substantial fraction of the total pool of money being invested, was there ever a time when "normal people" had a real say in where the money was being invested?
AFAIK, the only thing "normal people" can do is vote with their "feet" and pick a different prepackaged investment product, different investment company, or take their money and do the investment themselves.
Honestly, this comment of yours seems a non-sequitur and doesn't really address anything I said... you don't have the power to bend investment firms to your whim, but that doesn't mean that you need to -knowingly or not- carry water for the major LLM manufactures by perpetuating the "But think of how great the tools will be in the future!" meme. It has been years now, and everyone who has been paying attention can say with confidence that the LLM-based tools of the future are never that great... they're often not useless, but they're not worth the billions of dollars that have been and continue to be poured into their manufacture.
Related to your "We little people don't have any power anymore!" commentary, I note that TFA mentions that the kernel community has found that these LLM-based bug-finding tools have a false positive rate of ~50%. TFA goes on to mention that Coverity spent a huge number of years trying so hard to get people to buy its automated scanning software that had only a 20% false positive rate, and could not get enough people to buy the software.
Coverity went under because everyone hated how stupid and annoying the tooling was... at a 20% false positive rate. Once the hype machine starts slowing down, no one working at the coal face is going to buy a tool with a 50% false-positive rate. I personally very strongly believe that even if the tools had a 10% false positive rate, no one would pay the actual price that OpenAI and/or Anthropic would have to charge to recoup the research and manufacturing costs of a cutting-edge LLM-based bug-finding tool.
So no I have no way to address anything you said - that was the point, I don't believe you can - not with regulation and not with voting with your money (I am sure some people tried to vote with their money to slow down mega stores and keep the mom and pop shop alive - there maybe some of those still there, but largely it's big chains occupying most of the market) - that stuff hasn't worked - heck LLMs work better today that that stuff has ever.
I do think he repeats some myths in the video, or at least confidently states some things that are not demonstrated to be true, but his core point of 'you still gotta check these things' seems pretty solid.
I am genuinely curious what the myths/unproven things he states - I watched the video and it's repetitive sure but not much felt controversial to me.
"the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team."
https://docs.kernel.org/process/cve.html
Also as other replies said Linux kernel process is to assign CVE to everything - some of them may be just DDOSes, very hard to exploit and everything in between. All of them are bugs so they all get fixed and it's not a bad thing if distros ship those fixes and people update their kernel.
I'm fully aware of the philosophical arguments about transformative use etc. But I'm more interested in the concrete reality of high profile projects navigating this novel legal territory.
This is not a new problem. AI just changes the scale.
Mythos is amazing.
also, lol at "The bots are dumb - they want to please you line". LLMs have pretty much ruined technical collaboration between contributors. I get tilted every time an discussion has "but my claude said this..."
Do you mention this to disagree with the claim that the LLMs have been built to please their operator? If you do, I see no conflict between the claim that an LLM has been designed to please its operator and the claim that operators of LLMs tend to be absolute dogshit at considering statements from humans that conflict with claims made by that operator's LLM.
To rephrase my previous paragraph: It seems likely to me that an operator that has their ego repeatedly stroked by the output of the LLM they're using [0] will react very defensively when a human tells them things that disagree with the output of that LLM. "How dare you disagree with this thing that seems very human to me and consistently tells me things that I like!? Don't you understand how much I trust it because of how pleased it has made me?", yanno?
[0] ...thus, being very pleased by said output...
Wake me up when Raspberry Pis start refusing to open doors saying : "I'm sorry, Dave. I'm afraid I can't do that."
It’s all just pr stunts, fear spread fast and it’s very effective in marketing and spreading the word, which is effective, when I talk to some normal people they immediately bring the scary AI cyber attacks, kinda good as now all are willing to fund the industry!
He does not like the term "hallucinate" as it anthropomorphises a computer (pretends the computer is an "entity")
"NEVER upload any non-public information" - He's talking about how if you give Claude/GPT some secret info (like research, credentials, etc), it will train on it and give the same info to someone else. This is 100% the case for the free and consumer versions of these models, which is what most people use. For Enterprise plans they're not supposed to be doing this, but it's possible they will screw up and do it anyway.
The discourse has moved on for now, but 10-20 years ago when, why, and how to comment code was a hot topic. I'm sure the discourse will circle back, especially given how comments can be used to steer these models.
[0] ...which is apparently somehow involved in LLM manufacturing...
* «We've observed that these things have a 50% false positive rate. Coverity tried so hard but couldn't get people to buy their software, and it had a 20% false positive rate. No one is going to buy something with a 50% false positive rate.»
* «If you're going to use these tools, run them locally. The open-weights models that you can run locally are quite good enough. Anything you upload to the SAAS ones will be shared with other people... we've seen so many examples of it happening.»
In regards to the first point, I think he's failing to consider the fact that you can get most upper management to buy anything by providing them enough food, drugs, sex, and/or fear... but -otherwise-, yeah.
so is mythos just a chat bot with metasploit and its own cyber range?