Edit: And there was discussion about this back in 2024 as well
its-summertime 1 days ago [-]
For those with difficulty accessing:
- - -
From: Anthony Hurtado <[redacted since hn has no scrape protection]>
vpk_read_packet() divides vpk->last_block_size and (par->block_align - vpk->last_block_size) by par->ch_layout.nb_channels without checking for zero.
While vpk_read_header() validates nb_channels > 0, the codec parameters may become zero through format probing misidentification (VPK probe score is 2/3 of AVPROBE_SCORE_MAX) or codec parameter reset, causing SIGFPE.
Fix by:
- Checking nb_channels != 0 before division in vpk_read_packet
- Returning EOF for empty last blocks (last_block_size == 0)
- Validating block_count > 0 in vpk_read_header
- Validating last_block_size <= block_align in vpk_read_header
Found by fuzzing with libFuzzer + AddressSanitizer. Reproduces with
10 distinct inputs.
[patch redacted for brevity]
timpera 1 days ago [-]
Thank you! I gave up after more than 2 whole minutes of waiting on a high-end smartphone. I'm not sure this keeps bots out, but it definitely keeps users out…
1 days ago [-]
vachina 22 hours ago [-]
It keeps casual users (which most bots masquerade as) out. For frequent users of that site it is a solve once access forever.
iso1631 18 hours ago [-]
I don't get it, it briefly flashed up on my iphone , about quarter second maybe, then passed on. Does the same in private mode.
Maybe it uses other heuristics like detecting if someone lives in the AI agent world
duskdozer 16 hours ago [-]
Took like 2.5-3 minutes on my regular desktop browser. Nothing AI related
TheJoeMan 17 hours ago [-]
[dead]
jeroenhd 18 hours ago [-]
Damn, I've never seen Anubis set up that aggressively, I wonder what kind of attack their web servers must be under to set their bot filters up this strictly.
semiquaver 1 days ago [-]
Oddly enough I can’t access that site, it just heats up my phone solving hashes. Gave up after about a minute and anubis had only made it less than halfway through.
I doubt the real bots have any trouble bypassing it.
inventor7777 1 days ago [-]
It's puzzling how mild the reactions are to Anubis compared to the people reacting to seeing one singular Cloudflare captcha checkbox. I'd much rather a checkbox than a brief CPU-intensive hashing session.
jeroenhd 18 hours ago [-]
The proof-of-work approach is much better privacy-wise than what the usual CAPTCHA services are doing.
My ungrounded speakers have a tendency to pop and make noise when they come out of sleep (sleep? on a speaker? fuck you logitec) and every time these "simple" CAPTCHAs come up, even if I pass without solving their logic puzzles, I hear the speakers activate as the Javascript on the page is figuring out what kind of audio setup I have by playing a silent sound file.
The default Anubis config isn't really a problem for any devices I've tried, but the FFMPEG Anubis setup is quite extreme. I seem to be served the extra-difficult Javascript challenge, as well as a high-difficulty challenge, that takes even powerful computers quite a long time to complete.
Could just be countermeasures to the hug of death every website gets when they get linked on HN, though, but someone would need to set up auto-scaling for that.
tredre3 1 days ago [-]
Anubis is usually less obtrusive than that, though. This is the longest anubis challenge I've ever had, to the point of being absurd. Hopefully they have a genuine reason for having set the difficulty so high.
noir_lord 19 hours ago [-]
I wonder if it's a bug/issue with specific browsers.
It's near instant on desktop (Windows/FF/7950X3D) and I wouldn't expect the delta to be that large against a modern mobile device.
ailef 18 hours ago [-]
> I wonder if it's a bug/issue with specific browsers.
Seems like it. It loaded near instantly as well from my Android smartphone using Firefox.
zer00eyz 18 hours ago [-]
Safari takes minutes, FF is near instant on a M4 air.
And contrary to the directions, refreshing the page DOES help.
da_chicken 1 days ago [-]
I think when people are complaining about Captcha they're complaining about yet another "pick 6-20 pictures of traffic lights/school busses/stairs/stop signs/bicycles."
mapontosevenths 1 days ago [-]
If they want to train an AI they should pay for it like everyone else. Modern bots have zero problems solving these, it's just free training for them.
da_chicken 1 days ago [-]
> If they want to train an AI they should pay for it like everyone else.
They're paying for electricity and taking data without paying for it. It seems to me that they're paying for it exactly the same way everyone else in AI did.
Scoundreller 15 hours ago [-]
Does the bot include the rider when identifying bicycles? When is an ebike a motorcycle/no longer a bicycle? The support structure for a traffic light or just the coloured light bits?
inventor7777 1 days ago [-]
Not in this case. I wish I could find the actual post, but I recall reading a post on HN recently where a majority of the commenters were claiming that when they even see a Cloudflare verification checkbox that they leave the website.
This makes no sense to me as in my experience, you click the checkbox and then it verifies you without extra steps.
Izkata 24 hours ago [-]
Usually but not always. The challenge is after clicking the checkbox, if it can't manage to verify automatically. So those users have learned not to bother.
encody 16 hours ago [-]
For me it's very strange: I'd say about 7 times out of 10 it loads the checkbox for five to ten seconds, then I check it, then it loads for another five to ten seconds, refreshes the page, shows me a second checkbox, we go through the whole song and dance again, and then it lets me in.
anal_reactor 21 hours ago [-]
On my favorite browser (Opera Mobile with desktop mode) Cloudflare verification never works. It just says "failed, please try again" forever.
miki123211 21 hours ago [-]
[flagged]
xxs 21 hours ago [-]
the standard for user agents behavior tends to be described in certain RFCs
ShinyLeftPad 21 hours ago [-]
A zero-interaction screen is better. If I can open it in a new tab and then come back and it's fully loaded, it's good.
TeMPOraL 18 hours ago [-]
Unfortunately, those tools are there specifically to detect and block "zero-interaction activity".
ShinyLeftPad 18 hours ago [-]
I thought it just hashed something...
OroPla 18 hours ago [-]
Solving one captcha is mildly annoying. Being trapped in an infinite captcha loop will really grind your gears and eat away at your spirit.
Having my CPU go up for a while is nearly frictionless on the other hand. Worst case I'm stuck in a loop and the site isn't loaded when I get back to it, which is better than being stuck in a captcha loop and then not getting to the site.
Of course, not having to do any of that would be even better. I wish the concept of ZeroNet had caught on, where everything is hosted and served peer-to-peer. This gives you basically zero hosting costs and you are immune to DDOS.
59nadir 20 hours ago [-]
Cloudflare doesn't even let my browser (qutebrowser) through. Anubis will sometimes sit and ask for ridiculous amounts of work, but at least it's never outright denied access.
duskdozer 16 hours ago [-]
Well, Anubis actually lets me through eventually and most are very fast. Even this one is minimal compared to Cloudflare, which before I had to block its scripts entirely would just max out one CPU indefinitely (or at least a few hours, I found by accident, with no indication of stopping).
grishka 15 hours ago [-]
The Cloudflare captcha checkbox fingerprints the crap out of your browser. It's an opaque risk-based thing, which is the worst. Anubis just makes your browser brute force hashes, that's it.
blarg1 1 days ago [-]
> rather a checkbox than a brief CPU-intensive hashing session.
oh is that why my raspberry pi 5 can't browse websites anymore without freezing for a minute.
IanCal 20 hours ago [-]
I expect the complaints would be fewer if it was a smaller thing, or if everyone had their own version rather than it feeling like one company deciding if you should be able to use a large fraction of the internet.
hdgvhicv 21 hours ago [-]
Cloudflare doesn’t work most of the time for me. I’ve seen nothing on this site on my iPhone 12 mini.
TiredOfLife 19 hours ago [-]
Cloudflare checkboxes don't come with pictures of underage girls
zettabomb 19 hours ago [-]
Why do you have a problem with this?
TiredOfLife 18 hours ago [-]
I was explaining this part
> how mild the reactions are to Anubis compared to the people reacting to seeing one singular Cloudflare captcha checkbox
zettabomb 18 hours ago [-]
Ahh. My apologies, I seem to have interpreted your comment roughly opposite to how you intended it.
iso1631 18 hours ago [-]
proof of work requires my computer do do something and not me
cloudflare requires me to work for them
post-it 1 days ago [-]
It also took insanely long on my iPhone 16. "Made with heart in Canada" but configured poorly.
arjie 22 hours ago [-]
It's a pity about the web, because it's becoming less accessible over time. Regardless, I (and some others) archive many of the things we browse to, so here you go:
their anubis difficulty is wayyy to high; 6 is overkill
demibabs 1 days ago [-]
Yeah, is it trying to mine bitcoin or something? Anubis usually takes a second but here I waited a minute and got 20% through on a modern phone.
bulder 1 days ago [-]
Presumably they've configured it to use a higher difficulty challenge due to high rates of scraping on their bugtracker
myng111 1 days ago [-]
Difficulty 6 which some parts of FFmpeg use, is about the highest difficulty you can assign with the default Anubis config. For me personally I only serve that difficulty if I'm near certain the user is a bot. Serving it to everyone sure is a choice.
gguingff 1 days ago [-]
happy to report my bots have no trouble with anubis or any other pow mechanism, little bit of deno and i'm right through.
LoganDark 1 days ago [-]
The point is to deter bots that are scraping thousands to millions of websites in parallel, not user agents.
bartread 1 days ago [-]
Yeah, it's painful.
I get this crap when browsing on desktop a lot as well, principally because I stubbornly use Firefox as my main browser, and I habitually use a VPN when I connect my laptop to unsecured or even secured-but-accessible-to-large-numbers-of-people WiFi networks.
Like, seriously, bot detection "specialists", fuck off: I'm not a bot but your bot detection software IS shit, and I DO resent your shit software draining my battery and getting in my way. Learn to do your jobs properly, will you?
And don't come crying to me about how the problem you're trying to solve is "hard". I don't care: you chose it, you chose to considerably worsen the web browsing experience of millions of people globally, nobody made you. So go and find a different job if you're incapable of doing the one you have.
And if it's so "hard" why does your entire solution seem to be predicated on anyone's a bot if they're not running Chrome, or they are running an adblocker, or they appear to be from an unusual country that doesn't match their system language? Seriously, is this the level of sophistication you hacks operate at? To solve your "hard" problem?
You are extremely lame. Get out of my way.
Georgelemental 1 days ago [-]
> And don't come crying to me about how the problem you're trying to solve is "hard". I don't care: you chose it, you chose to considerably worsen the web browsing experience of millions of people globally, nobody made you.
Unfortunately, if you let all the bots in, they overwhelm your servers, and then nobody can access the website.
aystatic 1 days ago [-]
Not if you use a decentralized peer-to-peer Git forge like https://radicle.network. If one node goes down, users can still access the same issues/PRs from another endpoint.
breznev 22 hours ago [-]
Genuinely surprised to see these guys still committed to the grift. Berlin ain't so cheap these days, I guess
a2ff6eeb0 1 days ago [-]
I assume you're offering to pay for the increased server costs?
I had some git hosting up for a while, and was serving hundreds of qps and several terabytes per month. I can only imagine want significant sites are serving.
bartread 20 hours ago [-]
> I assume you're offering to pay for the increased server costs?
Such a non-argument.
I'm expecting people to create better, more effective, and less intrusive anti-bot measures. Measures that accurately detect bots but don't exclude real people from the web simply because of the browser they're using, or the country they either appear to be in or are in fact in, for example.
j16sdiz 20 hours ago [-]
I am expecting a unicorn.
bartread 19 hours ago [-]
It's weird to me that people are pushing back on me for expecting anti-bot services to actually solve the problem they already claim to solve.
a2ff6eeb0 16 hours ago [-]
Patches welcome.
8bitsrule 1 days ago [-]
It's indeed very annoying.
I've only seen 1 or 2 that know what they're doing. One's at lemmy.world ... just hovering over it is verified ...
1 days ago [-]
xxs 21 hours ago [-]
It took around 15kJ to access the site... that's a proper waste and somewhat sad, even though I understand.
pfdietz 18 hours ago [-]
So, about $0.001 of electricity. The nerve!
iso1631 18 hours ago [-]
A typical GPT5 or DeepSeek-R1 AI request uses about 100kJ
hexagonwin 21 hours ago [-]
it's putting full load on my 6core 11th gen i5 machine for more than a minute. i just closed it..
hiccuphippo 1 days ago [-]
Took less than a minute in my 5 year old xiaomi phone. It did take way longer than other Anubis sites I've seen.
yorwba 1 days ago [-]
A patch was submitted, but apparently not merged. That was also my experience trying to submit a patch for https://trac.ffmpeg.org/ticket/8738 . Somebody on the bug tracker took note, but was apparently unable to effect a merge in the intervening years.
Maybe now that ffmpeg is using Forgejo, the ball won't be dropped like this as often. Or there'll just be a five-digit number of open pull requests instead.
theowaway 1 days ago [-]
what the fuck is that anime catgirl bollocks
deepsun 23 hours ago [-]
It's Anubis and it's actually cool and loved project here. It's an open-source Captcha that filters out bots, and it doesn't track you around the web, unlike Google or cloudflare captcha.
stevekemp 21 hours ago [-]
Cool, and loved by some. Annoying and disliked by others.
I understand why people choose it, but if I see the catgirl I close the tab - same is I get the test from cloudflare.
deepsun 20 hours ago [-]
Yep, but it's their freedom -- website authors have freedom to designe websites however they choose, and we as consumers have freedom to not go there.
I was startled by the girl the first time as well, but once I learned what it is, I accepted it. Like a garden gnome on someone's front yard.
purerandomness 18 hours ago [-]
> Annoying and disliked by others
> if I see the catgirl I close the tab
Isn't it easier to just .. start loving it instead?
klez 21 hours ago [-]
> and loved project here
You may want to check upthread how loved it is :)
I, for one, don't hate it, but I hate what it represents and see its existence (rather, the reason for its existence) as a defeat for the web.
sunaookami 14 hours ago [-]
Speak for yourself.
phyzome 16 hours ago [-]
Huh, is this your first time seeing Anubis? It protects all sorts of sites now!
(By the way: Jackal, not cat.)
dabinat 1 days ago [-]
It’s interesting how AI may both raise and lower the quality of software. It’s very easy to send an AI agent on an open-ended bug hunt, and if it wastes a bunch of time and effort and finds nothing, no big deal. Time is much more important for a human developer with a salary.
dmix 1 days ago [-]
Finding the bugs with LLMs is easy. Reviewing the output, cleaning it up, and making sure it doesn't break something else is the hard part.
black_knight 1 days ago [-]
This is where I believe strong typing (like, Haskell-strong or stronger) and functional programming in general will be a win. The confidence I have that my fixes are localised when fixing Haskell code is infinitely stronger than fixing even Java, not speak about C, code.
astrange 1 days ago [-]
Haskell's type system would not easily prevent this bug. It's not good at numeric/logic issues like that. When people say "Haskell makes it impossible to write bugs" they mean "Haskell has enums" (ADTs).
_jackdk_ 1 days ago [-]
Liquid Haskell might require you to prove that the divisor is nonzero, but even in standard Haskell there's common idioms for ensuring that a list is non-empty (data NonEmpty a = a :| [a]) or that text is non-empty (newtype NonEmptyText = NonEmptyText Text, with non-exported constructor, helpers like make :: Text -> NonEmptyText, or more advanced tricks like https://exploring-better-ways.bellroy.com/haskell-koan-type-... ).
The big problem preventing this approach from working for numbers is that it's just so cumbersome there. Most of this is because all the arithmetic operators are bundled into a single Num typeclass, and `fromInteger :: Num a => Integer -> a` has a type that's impossible for a "non-zero number" wrapper to satisfy.
black_knight 1 days ago [-]
Definitely room for improvement on Haskell's standard library when it comes to the number-related type classes. Modern Haskell could do very well in this area with a good type-class redesign in this area. The issue I think is that this would invalidate a lot of existing code, relying upon that. But you can already replace Prelude with something else in your own code if you want to.
rootnod3 1 days ago [-]
I think Idris has a better chance there.
inigyou 1 days ago [-]
OOP has those too, and they're very annoying.
nh2 1 days ago [-]
In Haskell they are a little less annoying. It is just easier to reason about (including proving) pure functions.
inigyou 1 days ago [-]
I meant the constrained types by hiding the constructors. Super annoying, not automatically convertible, in Haskell you have to remember what the fake constructor is called, and write it every time you use it, but at least it's efficiently implemented with newtype, unlike the Java OOP version. Think about writing a value with several nested constrained types, like NonEmptyListOne (makeNonZeroNumber 42, 'h' `NonEmptyString` "ello world"). It's just really annoying.
_jackdk_ 1 days ago [-]
The blog link I mentioned avoids this cost with literals, by providing using a required type argument to check the string length at compile time without TH. It requires a relatively recent GHC:
make :: forall symbol -> (IsNonEmptySymbol symbol) => NonEmptyText
type family IsNonEmptySymbol symbol :: Constraint where
IsNonEmptySymbol "" = Unsatisfiable (Text "Expected a non-empty string")
IsNonEmptySymbol _ = (()::Constraint) -- empty constraint is always satisfied
black_knight 1 days ago [-]
I am not claiming you cant write buggy code in Haskell! But following good functional style, your bug will more likely be compartmentalised, and fixing it will not break some other part of your program.
StilesCrisis 1 days ago [-]
You can write good functional code in many languages. (Even C++!)
black_knight 1 days ago [-]
Sure! I have done my fair share of pretending Java and C++ support my functional style. But at the end of the day, you have better support for writing that style in a real functional programming language. And I wonder how well one can enforce a functional style in say Java or C++ upon the LLMs. Who knows, they might be great at it?
tome 22 hours ago [-]
People don’t say "Haskell makes it impossible to write bugs"! You may have heard "if it compiles it works" which is somewhat tongue in cheek, but also true for a sufficiently loose interpretation of "works" in a way it is not true for languages with a less strong and flexible type system.
deepsun 23 hours ago [-]
You haven't mentioned the dynamic typed languages that I believe should die -- Python and Javascript. The only good use case for dynamic typing is notebooks (niche of R lang) where you're throwing out the code you just wrote after getting the result you wanted from it.
theLiminator 1 days ago [-]
Imo, formal methods like more expressive/stricter type systems are key to making LLM generated code successful. Of course models will get better, but trusting the output will become much easier with a type system that proves more properties.
fouronnes3 1 days ago [-]
What's stronger than Haskell?
black_knight 1 days ago [-]
Dependent types is one possible direction. Not sure when a language with dependent types will arise which will be useful for making real programs.
Agda is the most mature dependently typed programming languae (having been around since the 90s – it is basically Haskell on steroids), but has a more proof-assistant flavor than an actual programming language flavor. Opus & Fable write Agda quite well, so LLMs can understand dependent types.
astrange 1 days ago [-]
Anything with ranged numeric types. Like everyone's favorite functional programming language, Ada.
ghaslt 1 days ago [-]
This issue raises SIGFPE. Ada would raise Constraint_error, which is easier to catch than a signal, but still occurs at runtime.
You need range proofs to be 100% safe, and then you can as well use the regular type because invalid values will not occur.
black_knight 1 days ago [-]
Or Liquid Haskell.
TheGoddessInari 1 days ago [-]
Lean 4, Idris 2.
theLiminator 1 days ago [-]
Perhaps coq/agda/idris/etc.
UltraSane 1 days ago [-]
Even Lean 4 strong typing
sadfgknerknksdf 1 days ago [-]
If finding the bugs with LLMs is easy. Then making sure it doesn't break something else is just LLMs finding no bugs. Easy.
BikiniPrince 1 days ago [-]
That hasn’t been that bad. My real issue has been the time sink involved in following along with the maintainer and jumper through their hoops. Even after I demonstrate a flaw and a potential fix. My schedule is just so busy I need to pencil in time to deal with them.
hombre_fatal 1 days ago [-]
The missing part of this is that verifying the bug with LLMs is also easy, and so is adversarially reviewing the proposed fix with LLMs.
The only thing left for you to do should be directional decisions. The LLMs should pause and rope you in if the fix involves directional/invariant changes.
1 days ago [-]
nonethewiser 1 days ago [-]
No one can keep up with the volume of code AI produces.
We wont stop using AI.
We will use AI to check AI.
Of course this is crazy, but it will also unlock pretty insane scaling and productivity and ultimately we will manage it on either end via requirements and tests.
adamddev1 1 days ago [-]
> it will also unlock pretty insane scaling and productivity
Insane scaling of bloat, bugs, and technical debt I'd say.
> We will manage it on either end via requirements and tests
It is so crazy that this is being touted as a sane strategy. When I was a much worse programmer, I tried to write a big complicated string manipulation function to take two types of scripts in a language and add diacritics. I had the requirements very clear. I had the tests very clearly with all the edge cases. But I didn't have a good and clear picture of how to attack the problem which was quite novel for me. As I got closer to passing all the tests it got exponentially more unruly and confusing. And nearing the end I was frantically changing little bits here and there wincing and praying and hoping the tests would pass. "Please work! Come on!" Then when I got close enough, I could never ever think about touching that mess again.
I was a below average programmer then throwing myself at some novel problem I didn't understand. Throwing LLMs that produce below average code at novel problems and relying on tests and requirements is not where we want to go to make real progress.
(Years later after much learning and coding myself I was able to redo the function in a totally different way. This time I actually understood how to attack the strange problem and made something clean, clear, and robust that just worked. The tests then become a secondary guardrail, not the main force of correction.)
We are seeing such a massive regression from what we've learned over the years of CS.
shiandow 1 days ago [-]
I think all code is technical debt in a way. Good code is a necessary evil, bad code is more evil than necessary.
Generating code automatically when you're not even quite sure what it is or even should be doing is insanity.
nextaccountic 1 days ago [-]
I'm not so sure LLM code today is below average. There was a time that things posted to dailywtf were normal everyday stuff
krupan 1 days ago [-]
Sorry, no, they wouldn't have been WTF's if they were normal
You shared a story of a novice incompetent human programmer and this should tell us that AI is bad at coding.
nonethewiser 1 days ago [-]
>Insane scaling of bloat, bugs, and technical debt I'd say.
You just described every legacy codebase. Many of which are widely used and do a lot of sales. You dont need a clean codebase to have a valuable product.
>It is so crazy that this is being touted as a sane strategy.
Re-read what I said. I literally called it crazy.
It is the same dynamic that gave us customer service from some call center in India. Why would companies do this? Customer service got worse. Are they stupid? No, it's just worth it. The quality goes down but the business can scale more so it doesnt matter.
AI will absolutely be good enough at doing things that we'll happily accept some jankiness at times so that we can devote an extra 3000 hours per year per person to other things.
Im not even suggesting its a good thing. I just think the incentive structure dictates it. You're not going to have time to maintain a small slice of some service by hand.
harambae 1 days ago [-]
It's mostly (not entirely, but mostly) finding security issues in old human-written code. It'll eventually start running out of those.
From that standpoint, it's not a crazy setup security-wise. Maybe still crazy for development.
stefan_ 1 days ago [-]
You can point AI at any AI produced code and ask it to review it, get back 10 bullet points and a few pages of prose. And the fun part is, you can do that over and over and over again!
bilalq 1 days ago [-]
This happens all the time. Yesterday, I ran into an especially egregious case.
I had Fable add a new subcommand to our internal CLI tool. I reviewed and tested it locally and had to suggest several fixes that I feel like I wouldn't have had to tell a human senior engineer to do. When it finally submitted the PR, I had it on a loop waiting a few minutes for comments on the PR, then assessing/addressing/replying-to/resolving them, and then repeating again until all AI reviewers were okay with it. It ended up going through dozens of revisions and ended up with 160 comments left on the PR.
krona 1 days ago [-]
You're suggesting that LLMs get better at fixing bugs/vulnerabilities, but at the same time stop getting better at finding them? What if this difference is inherent and essential?
TacticalCoder 1 days ago [-]
> You're suggesting that LLMs get better at fixing bugs/vulnerabilities, but at the same time stop getting better at finding them?
Are you implying that all code writing by LLMs atm is bug-free?
krona 1 days ago [-]
Absolutely not. By most accounts they're terrible at fixing anything other than trivial bugs in complex codebases e.g. Linux kernel, but they're much better at finding them.
a2ff6eeb0 1 days ago [-]
So you put it in a loop and tell it to find the bugs in the code it wrote. What's the issue?
krona 23 hours ago [-]
This is a self confession if I ever saw one.
a2ff6eeb0 16 hours ago [-]
I absolutely do this. It works great.
Planktonne 20 hours ago [-]
What value are you providing in this scenario?
a2ff6eeb0 16 hours ago [-]
Manual testing, and making sure that the AI didn't create so many bugs.
But, to the underlying question, obviously as we automate more and more of our work, of course we provide less and less value. We're heading towards a future where selling thought for money isn't going to work so well.
kayamon 1 days ago [-]
Volume..... <sigh>
It used to be considered a quality of good code that there would be less code, not more.
Some people always tryin to get the highscore on golf.
adrianN 1 days ago [-]
You can have both less code per problem and more code overall when you make problem solving cheap enough.
CPLX 1 days ago [-]
In fairness at root this has been going on for awhile. No one can keep up with the volume of machine code that modern more abstracted codebases produce.
We didn't stop using syntactic programming languages we used code to check code.
Not sure it's really crazy at all. It's been an abstraction for programmers probably since we stopped soldering transistors to each other.
ldng 1 days ago [-]
There is a MAJOR difference between predictable generated machine code and Russian Roulette code generator.
CPLX 1 days ago [-]
Of course there is.
But if you don’t actually read it…
1 days ago [-]
macless 1 days ago [-]
[flagged]
1 days ago [-]
bewareofscams 1 days ago [-]
[flagged]
shevy-java 1 days ago [-]
> LLMs do find bugs, do save time
They find bugs but whether they save time is nowhere near as clear as you try to insinuate here.
pixl97 1 days ago [-]
They save time in finding bugs.
lukan 1 days ago [-]
And for me they also save time in fixing bugs.
owebmaster 1 days ago [-]
Unless it's finding a bug it added then it's time wasted x2
simonjuk 1 days ago [-]
In my experience, there are two ways to use AI: speed or quality. Speed is where you give the AI a task to do and you review it; quality is where you write the code yourself and you get AI to review it. Both are valid for different situations.
merb 1 days ago [-]
My plan for bigger things is mostly:
Generate multiple solutions- they do not to work 100% correctly.
And than I check which I would prefer. Which is more to our applications taste.
And than I would take the vibe output as a kind of a ‚plan‘ which I use to implement but not follow 100% and at the end I take my solution and review it.
I gain speed with that because I often can quickly see the pros and cons of a solution way better than when I would manually do it and hang on a major roadblock and also I even see such roadblocks in the vibe output - it’s mostly the part with an unnecessary amount of new code that looks nonsensical.
UltraSane 1 days ago [-]
Using a LLM whose output is slowed to the rate of a human programmer as a pair programming partner is a very interesting experience.
evenhash 1 days ago [-]
> It’s very easy to send an AI agent on an open-ended bug hunt, and if it wastes a bunch of time and effort and finds nothing, no big deal.
No big deal? It’s not like it’s free… tokens cost money.
rogerrogerr 1 days ago [-]
Often rounds to free compared to human costs.
1 days ago [-]
UltraSane 1 days ago [-]
When talking about LLM tokens the cost is almost always being implicitly compared to very expensive human developer time.
DarmokTanagra 23 hours ago [-]
Having worked in a few vibe coded codebases over the last few years I can safely say that AI is not raising the quality of anything.
tikotus 21 hours ago [-]
I had the same knee-jerk reaction. "Did I read that correctly?"
But yeah, I guess it can be used to increase certain aspects of quality by letting them go wild. But I think I mostly hear about security or crash issues. In my experience they don't outweigh the number of other issues they cause. Like UI bugs. I've seen more than one service constantly rolling out features that are completely broken, just to have a completely new, still broken, solution available the next day.
Supermancho 1 days ago [-]
I don't care if you call it an over-engineered looping machine or what, there are concrete benefits to using LLMs for this. They work faster than developing your own looping algorithm and more often produce useful results than not.
saghm 1 days ago [-]
It's not even like fuzzers are valuable because of the process they use specifically either; the value is that they produce a concrete input that you can use as a reproducible test case at that point. The value could be produced by gazing into a crystal ball for all I care, as long as I can use what it gives me to reproduce a bug.
eviks 1 days ago [-]
But what's your expectation of the net?
shevy-java 1 days ago [-]
I dislike AI, but if AI finds real bugs then this is in my opinion objectively a positive thing. Of course the question is what constitutes a real bug.
1 days ago [-]
pixl97 1 days ago [-]
Unfiltered models will help build exploits for the bugs they find, so there is some means of measuring their efficacy.
klipt 1 days ago [-]
If you're just talking about security bugs.
There are also non security bugs that don't have exploits but just make the user experience worse.
hn_submit 1 days ago [-]
A.I. is useful for this. But it would be even more useful if all new code were written in Rust or some other memory-safe language.
A.I. could also be used to port C/C++ codebases to Rust, which isn't economically feasible at the moment.
senderista 1 days ago [-]
AI will have plenty of security bugs left to find in Rust codebases.
Spivak 1 days ago [-]
I mean I get the sentiment but Rust won't save you against division by zero, it'll just panic at runtime like every other language.
Gigachad 1 days ago [-]
From a security perspective, panic at runtime is not that bad for security. Much better than continuing to run with undefined behavior. If someone sends a malformed video in and it crashes the ffmpeg process you can just log it and restart it. Vs potentially exploiting the system.
Sharlin 20 hours ago [-]
The Rust standard library has `NonZero<T>`, which, if used, at least forces you to consider what you initialize it with. Doing
let foo = NonZero::new(unvalidated_input).unwrap();
is at the very least a big red sign that stands out in the code and should fail code review.
skupig 1 days ago [-]
Am I missing something? Who cares? This isn't a security issue, it's just an unexploitable crash on bad data.
inigyou 1 days ago [-]
No, you're not. It's a minor bug, probably with an easy fix, that deserves to be fixed. It's not worthy of front page HN...
Jaxan 23 hours ago [-]
I guess it’s submitted for the method rather than the result.
mcdow 15 hours ago [-]
that’s the issue with these newer models. they are able to string together a sequence of “not-serious” bugs in a system that ultimately results in some serious vulnerabilities.
it may not be an issue for ffmpeg, but it might be for an application that bundles ffmpeg.
ramon156 21 hours ago [-]
everyone knows HN only accepts security write-ups /s
ChannelFence 14 hours ago [-]
Two months and 1100+ commits to rediscover a bug that was already found in 2024 is probably the funniest possible ending to a "vibecoded fuzzer" story.
13 hours ago [-]
dclavijo 14 hours ago [-]
No, the funny thing is that this comment is comming from an account with 1 karma, no submisions and only 1 comment, I wonder if this is your only account and you just came to spill hate or you are using multiple accounts to discredit other peoples work.
qwt1254 14 hours ago [-]
What work? You let a plagiarism machine write something and it didn't produce any new bugs.
dclavijo 13 hours ago [-]
another sockpuppet account, is funny how they come so fast
ChannelFence 14 hours ago [-]
4 now lol. why would my karma matter anyway?
this isnt reddit
dclavijo 14 hours ago [-]
It just states your intention, you could just have been constructive with your first comment on HN.
ChannelFence 14 hours ago [-]
Why is it a bad thing I found this funny?
Two months and 1.1k+ commits to rediscover an already known bug IS funny. That doesn't mean I hate the person or want to discredit their work. You're reading way too much into a comment :)
dclavijo 14 hours ago [-]
Humor is subjective, not objective; if we delve into the subjective realm, I can feel whatever I want, just like you. You haven't offered anything constructive yet. I'm more willing to listen to your ideas if you have any.
ChannelFence 13 hours ago [-]
I really wasn't trying to offer anything constructive. I was just pointing out something I found funny and made a joke about it. It wasn't anything more serious than that.
qwt1254 14 hours ago [-]
The comment was constructive. It exposes yet another AI lie!
fleroviumna 14 hours ago [-]
[dead]
ks2048 1 days ago [-]
No doubt fuzzers (vibecoded or otherwise) can be powerful, but can't you just mark all "/" as potential divide by zero errors?
I guess sometimes developers think they "know" some variable won't be zero, but unless it checked explicitly or by the compiler, that shouldn't be trusted.
Someone 1 days ago [-]
> but can't you just mark all "/" as potential divide by zero errors?
If you’re accepting large false positives rates: yes.
If you want users to take your warnings serious: no.
(Nitpick: you certainly don’t want to flag _all_ of them. Divisions by non-zero constants definitely should be excluded, for example (integer division by -1 can lead to overflow, but that would be a different warning))
saghm 1 days ago [-]
Fuzzers find inputs, not just "potential" errors that aren't triggerable.
dooglius 1 days ago [-]
What are you suggesting and how would it be different than how SIGFPE already works?
MaxBarraclough 1 days ago [-]
If it's possible for program execution with some particular input to lead to a divide-by-zero, that's a bug, especially if the program is expected to be able to handle malformed inputs, or perhaps even deliberately malicious ones. It's not trivial to determine whether a program does this correctly. If it was, program analysis would be easy.
Division can 'go wrong' for certain inputs, but it's not just division. In C, signed integer addition, subtraction, and multiplication, all give undefined behaviour on overflow.
As 'Someone' already pointed out, it's not helpful to just flag all uses of the division operator, or of other potentially dangerous operators. Minimising false positives is one of the core challenges of program analysis.
wvbdmp 1 days ago [-]
I mean there could be a guard clause? But yeah, seems like this could be statically evaluated like how some IDEs see a null check and don’t complain about nullability within the same scope.
throwa356262 20 hours ago [-]
I am sure the fuzzer is interesting.
But this bug feels like something an LLM would flag as a major finding but turns out to be completely benign.
Update: I tried to look into the fuzzer but it is hard to get past the AI blabb. Can someone please explain to me what it does beside being structure aware?
boomlinde 20 hours ago [-]
It crashes because of input that should have been rejected for being invalid. How could that be construed as being benign?
throwa356262 19 hours ago [-]
Because an attacker would not gain anything he not already has. This is basically local self-DOS.
boomlinde 19 hours ago [-]
That's not a quality of ffmpeg or this bug, but of the application you use it for. If you only expose your ffmpeg-based application to your own input then yes, of course it's a self-DOS. But if you, say, expose it as a web service passing arbitrary user input to ffmpeg, that no longer holds.
throwa356262 18 hours ago [-]
Even then it will be a self-dos: the video you uploaded won't be processed.
boomlinde 17 hours ago [-]
Again, this is a crash bug, and again, whether it's a "self-dos" isn't a quality of the bug or ffmpeg.
The implications of the crash depends entirely on the implementation of the process it crashes. If I use ffmpeg as a library it'll crash my process upon processing the offending file. How is my process designed? How is every process that uses ffmpeg designed? You don't know, therefore you can't say that it's a "self-dos" in every case even if you know that it is in some cases.
Maybe I am clever enough to have read up on the history of ffmpeg vulnerabilities before deployment to an attacker-facing service and have designed a solution where a crash has minimal implications, but maybe I'm not, and haven't. It's beside the point.
j16sdiz 19 hours ago [-]
afaict, decoder bugs like these are treated with lowest priority possible.
It is not enabled by default. It is used only in video games, which input files are fixed set of asset that came with the game.
It can be a crash, yes. but the typical user of this codec won't care.
cptroot 1 days ago [-]
This is not a real bug in FFmpeg. This is a demonstration that if you control a custom AVIO module it is possible to crash FFmpeg by giving it bad data.
inigyou 1 days ago [-]
Not custom. It's an existing module for a format called VPK. It's a quite trivial bug though, not exploitable apart from DOS and won't ever happen in a real file.
VladVladikoff 1 days ago [-]
I even question if it is a DOS vector. So the thread crashes and then the system that controls the threads cleans it up and opens a new thread. Seems to be a trivial impact, unless it locks up the thread somehow.
inigyou 1 days ago [-]
Threads don't work that way. A fatal exception on any thread kills the process.
VladVladikoff 1 days ago [-]
And the parent will spawn a new process. Unless the server is terribly poorly misconfigured.
Edit; for what it’s worth I’ve run a server processing video with FFMPEG for 10 years now, and there’s just so many things that can make FFMPEG crash. All sorts of corrupted videos people upload. If your server doesn’t recovery gracefully from a crashed FFMPEG thread, that’s on you, not FFMPEG.
LoganDark 1 days ago [-]
I thought you meant Disk Operating System until I realized you probably meant DoS
avadodin 20 hours ago [-]
FFmpeg on DOS is enough for anybody as long as you let your 0.00066B model check the movie for 0day exploits.
justonenote 1 days ago [-]
Whatever about the specifics of this bug and whether its a useful vector, this is not surprising even in the slightest?
My current opinion on LLMs is that they are superhuman in that they lack fatigue, they have close to full knowledge across all subjects which are known to humans at least publicly, and the fact that you can vibe code a harness to look for bugs in a famously complicated C codebase is intern level stuff and hardly news.
Smart aspiring blackhats will be targeting tmux next, both with light llm jailbreaks, light supply chain attacks (web search results) and LPEs within certain environments which weren't particularly useful before but with agents running on auto mode for hours become a very valuable springboard. I'm not sure on the quality of tmux code but I know its written in C and is very complex and was not at all designed to defend against this type of threat.
hnlmorg 1 days ago [-]
I don’t think tmux is the most worthwhile target because you’d need the user to either execute code locally (thus negating any point in targeting tmux) or rely on the user curl or cat some compromised document (in which case you’re better off targeting curl or cat).
justonenote 1 days ago [-]
the point is tmux is being used by many developers working in high value targets to automate long running unsupervised agent tasks. you don't need the user to execute code, you need _their agent_ to stumble on the wrong search result or github repo and it wont be noticed for hours that they loaded a persistent threat into your environment.
hnlmorg 1 days ago [-]
That seems even harder to do because an agent wouldnt be output text verbatim, which means you cant make use of a rendering bug (eg parsing escape codes).
So you’re back to depending on the agent to execute code locally. at which point you’ve already compromised the system so don’t need a tmux bug.
I’ve spent a lot of time in tmux. Including writing a frontend for it. So I’m probably more familiar than most. And I hear a lot of people say tmux (specifically) is a vulnerability because it’s written in C. But I struggle to see how it’s any more of a vulnerability than (for example) coreutils. Or any other piece of software for that matter.
jonhohle 1 days ago [-]
Not that it doesn’t have issues, but I’m not sure why you’d choose tmux of all things. It runs as a user and has no privileges to escalate. It was written for and is part of OpenBSD and follows their security hardening practices.
(There actually was one privilege escalation bug in tmux, but it actually seems like a distro packaging error. The distro setgid the executable so the resulting shell inherited the additional group. This didn’t require any exploit, that’s just how child process inheritance works.)
justonenote 1 days ago [-]
as I mentioned in another sibling, its because it's a very common denominator in high value targets. I didn't know its legacy was from OpenBSD but I really doubt that that helps it much in this scenario, when I say LPE I'm not talking about user to root elevation, I'm talking parsed text/control sequences to arb code execution in the user context. These will slip past llm classifiers as safe and I'm fairly sure that they are extremely common in codebases like tmux, despite them having strong security posture its just a threat that was previously a bit outlandish and not accounted for.
persisted malicious code running in your tmux process that you don't know about is probably not where you want to be, for obvious reasons.
hnlmorg 1 days ago [-]
Agent harnesses aren’t going to output ansi escape sequences verbatim to the terminal.
If you wanted to booby trap a repository then you’re far better off with a prompt injection attack.
1 days ago [-]
1 days ago [-]
senordevnyc 1 days ago [-]
the fact that you can vibe code a harness to look for bugs in a famously complicated C codebase is intern level stuff and hardly news
It seems like this would have been pure fantasy not that long ago though. So why isn’t it noteworthy again? I don’t really follow what you’re complaining about.
1 days ago [-]
soiax 17 hours ago [-]
Why are people upvoting a unexploitable bug? How is this interesting? There are thounds of these, no one even reports them unless they are exploitable, DoS only.
21wa 16 hours ago [-]
Because the fuzzer was vibe coded and stolen by Claude! They only need the headline for celebrating another "AI victory" on Twitter, even though the issue was found in 2024 by OSSFuzz:
It is interesting that FFmpeg has its own Git server. Maybe we should move there too?
snailmailman 1 days ago [-]
Lots of projects run their own git or forgejo or similar. I run my own private forge, and it has a higher uptime than GitHub. (A shockingly low bar, tbh)
It’s surprisingly simple to setup, and the hardware requirements are pretty small for a private or small forge, as it’s usually a relatively small number of users/repos/etc.
sva_ 1 days ago [-]
You can add several remotes to your git, and I'd recommend you do so.
TacticalCoder 1 days ago [-]
> It is interesting that FFmpeg has its own Git server. Maybe we should move there too?
Git is a DVCS. I know many people only ever used Git through Github and forgot what the 'D' in DVCS means but whether or not they remember what the 'D' stands for, running your own Git server is trivial. Especially in this day and age of LLMs were you can just ask: "Clone this repo and convert it to base Git repo and serve it on the LAN PLZ KTHX".
The result is going to be more stable than Github and, arguably, more secure too.
inigyou 1 days ago [-]
If you have SSH access to a server and Git is installed on that server, you can use it as a Git server. No additional setup is required. The Git client knows how to log in and invoke the Git server over SSH.
tensegrist 1 days ago [-]
note that this seems to be a bug in what i expect (feel free to correct me) is a code path for a little-used codec
it's widely used but in "industry" applications. so ffmpeg is probably being used in a lot of offices (studios) and maybe even being included in end user software.
inigyou 1 days ago [-]
Understatement of the year. Almost everything that processes video uses ffmpeg.
BikiniPrince 1 days ago [-]
Funny thing, I know I'm brushing up against something in gStreamer developer, but Fable flips out. I have only a loose idea where the issue might be lurking.
Next week, I'll apply for the cyber and I suspect I'll find something similar.
Right now, it's just annoying and thanks the OpenAI cyber was much easier to get access to.
written-beyond 19 hours ago [-]
IDK seems like a bug that could've taken a human a few minutes at best to find. I found a bug in SystemD that would crash the daemon because a bad SystemD unit file configuration. That took me like 5 minutes to actually track down in the actual source code.
I understand the utility of this though, I just don't see this particular bug and something that would be particularly difficult o find pre LLM era.
thephyber 17 hours ago [-]
This is the wrong mentality.
The fuzzer found the bug before any humans did, so there is a mismatch of developers who could find this bug and those who did (without an LLM-coded fuzzer).
The value of the fuzzer continues long after it found this one bug.
It's worth nothing that in the bug discussion thread, the bug fix author pointed out that it's not easy to set up the config then call the functions in the right order. Your comment assumes that the reader has enough context to read the code and build the finite state automata in their mind. The bug fix reporter's comments suggest that you are assuming things which you shouldn't assume.
written-beyond 7 hours ago [-]
But I am being very specific to this use case, where a division by zero bug was found. Why couldn't you have just grepped through the codebase, found all possible divisions and ensured that they had a check on it to never be less than or equal to 0?
The bug I located in SystemD was literally a null ptr exception. All they had to do was perform a null check on a cstring but they hadn't.
I don't see the utility of reporting an LLM made fuzzer finding bugs that could be found by a lint rule or static analysis. I will appreciate a post about an LLM fuzzing software to find unique corner cases, which I predict will happen soon, in ACL controlled systems caused by policy shadowing.
pfdietz 16 hours ago [-]
Bugs are always easier to find in retrospect.
written-beyond 7 hours ago [-]
This is literally the lowest rung of a bug, it's equivalent to finding a nullptr exception. You can literally avoid them with 1 if statement.
OP here: You are welcome to send a PR if you like. I'll be grateful if someone makes the readme more human.
speps 21 hours ago [-]
Surely, you're the best person to do that..? Unless you're not human of course.
Zebfross 1 days ago [-]
Why submit an issue rather than just making the fix and adding the tests in PR? Seems like they're just making work for the maintainers.
dclavijo 1 days ago [-]
OP here: A bug report just needs a proof of existence for the condition while a bug fix needs a proof of correctness. Sometimes is the best to let the developers who are day to day in the codebase to choose the best fix and if they what to fix it.
dclavijo 18 hours ago [-]
OP here: for those interested is not that an LLM found the bugs the fuzzer found them.
My take on this matter was to implement as much information theory algorithms as possible, and try to extract as much structure with statistical importance from the binary being fuzzed. Also port as much features from other fuzzers and whie papers on the mater (llms are good at connecting dots across vast codebases and papers).
I honestly can not take full credit for this work since I made it with AI, but I has taken two months of my time and 1100+ commits.
My developing process was to use several models from several vendors not just Claude that decouples it from a single vendor/model and throws to the flor that llms regurgitate verbatim code.
Also the interesting part is the developing pipeline I have had setup my own cicd with my own tool impactguard whitch saved me a couple of times and hard rules on the agent.md(70% to 80% of those rules i wrote them by hand).
The pytest testing battery is also interesting, I adopted TDD and to me since I adopted it seems that llms make less bugs.
Yes I know the code and the readme might look like ai slop as pointed out earlier but is efective at finding bugs. At the end of the day is all economy: you spend a lot of tokens once and keep the fuzzer forever, not the same as paying every time for tokens to find bugs.
As pointed out in the readme, this fuzzer trades speed for edge novelty, maybe there is it's niche.
Also we found earlier another bug with this fuzzer https://code.ffmpeg.org/FFmpeg/FFmpeg/issues/23945.
For the concerned IMO: rather than the results the methodology is more important.
I welcome constructive criticism and feedback. Any input is useful to me.
nixpulvis 14 hours ago [-]
I can't speak to the quality of the fuzzer since I haven't used it or looked at it thoroughly, but it does seem to cover a lot of ground on features and interesting concepts. I'll definitely be reading more into what you have here.
peter_retief 22 hours ago [-]
Bugs days are numbered with AI!
pfdietz 16 hours ago [-]
Technically correct, because there are infinitely many numbers.
1saadcodes 1 days ago [-]
I find it pretty cool that a fuzzer thrown together this way actually found a bug in ffmpeg
pfdietz 16 hours ago [-]
The thing about testing is that each time you produce a new kind of tester you have a chance to find bugs in the blind spots of the previous testing approaches. Diversity makes sense, more so than in software construction.
dclavijo 16 hours ago [-]
Yeah that was my bet, escaping the local minima imposed by the current state of fuzzing. I am seeing this problem as statistical and information theory problem.
Some day someone by chance will create another fuzzer that would find more bugs because of the blindspots in the previous generation including mine.
krpovmu 18 hours ago [-]
Did we find it?, or Did AI do my job?
thephyber 17 hours ago [-]
Do you normally get paid to fix bugs in open source repos?
sylware 19 hours ago [-]
The real core of the issue is actually the complexity/size and core design of media container/codec file formats.
cpriest 1 days ago [-]
Nice find. The interesting part isn't "AI wrote the fuzzer." It's that a cheap random harness still hits classical bugs in ancient parsers. Keep the corpus; throw away the hype.
robertlagrant 1 days ago [-]
What we need is a numeric type that cannot be zero.
winwang 1 days ago [-]
Every day, we stray closer to Haskell. Dare I say it: good!
drdaeman 1 days ago [-]
What we need are refinement types, where there’s a base type and a predicate. F* has this:
val (/) : int -> (divisor:int { divisor <> 0 }) -> int
yeputons 1 days ago [-]
And also cannot be INT_MIN, otherwise -1 / INT_MIN is undefined behaviour(!) in C and C++.
roadbuster 1 days ago [-]
The only way to achieve this is to either put a runtime software check on a variable whenever it's assigned/used, or to literally add hardware support in processors themselves which literally throws an interrupt when a "neverShallBeZero" variable is assigned to zero.
There's no viable way to statically prove at compile-time that these variables will never become zero at runtime, ultimately forcing a system of endless runtime checks (be it software or hardware)... which is why processors already throw exception interrupts when division by zero is attempted.
inigyou 1 days ago [-]
It's possible, just extremely difficult.
colechristensen 1 days ago [-]
You're kind of saying the only way to do it is in software or hardware :)
The projectively extended real line defines division by zero, no reason you couldn't have a floating point type that implemented it.
>There's no viable way to statically prove at compile-time that these variables will never become zero at runtime
strongly typed programming languages like Ada allow for types which have ranges such as disallowing zero -- but also any arbitrary thing like you can create a floating point "degrees" type which is [0.0, 360.0] or any other ranged type
rhdunn 1 days ago [-]
It would be more flexible for a compiler to reuse the range analysis logic used in optimizations for statically verifiable divide by zeros. That way you could extend it to other things like statically verifiable overflows.
duped 1 days ago [-]
For stuff like niche value optimization sure. For practical arithmetic code, nah. Like with this bug, all that changed is that garbage data in gives the user an error that they tried to process garbage data. Adding a new type doesn't make the code better, it just moves the error around. And you really don't want an infix division operator to fail to type check if the right hand side isn't a nonzero type, do you?
jeffbee 1 days ago [-]
I imagine the discussion will center around this application of AI, but to me this is just the Nth proof of the proven fact that you must build ffmpeg, if you insist on using it, with only an allow-list of file formats that you expect to encounter, and not with the kitchen sink of stuff you are never going to need.
dorianmariewo 21 hours ago [-]
given enough ai, all bugs are shallow
thephyber 17 hours ago [-]
Given enough AI, all bug fixes have extremely deep carbon footprint.
Surac 1 days ago [-]
send patches
rs_rs_rs_rs_rs 1 days ago [-]
...they did.
ligarota 1 days ago [-]
Where?
They only suggested a basic guard, chich can be useless if this case never happens
12j3afAv 1 days ago [-]
Generating an incorrect input file seems to be the easiest task of all for any fuzzer.
Generating correct input to get deep into the call stack and then finding something is the hard part.
1 days ago [-]
ozereray1 16 hours ago [-]
[flagged]
aaron695 1 days ago [-]
[dead]
akshay_akula 1 days ago [-]
[flagged]
wy35 1 days ago [-]
Unrelated to the submitted link -- just checked your comment history and all of your comments are AI-generated like this one. What's the motivation for this?
f311a 1 days ago [-]
He won’t reply, he’s busy promoting himself and his peojects with AI.
akshay_akula 16 hours ago [-]
Also haven't done any self promo, but yeah i'll stop running my posts thru chat
bigfishrunning 1 days ago [-]
probably karma farming
akshay_akula 1 days ago [-]
LoL I just sound botted
whatsThisBtn4 1 days ago [-]
[flagged]
VCFundedGenYer 1 days ago [-]
The fruits of using LLMs to code.
You'll waste far more time finding what it quietly and subtly wrecked than you would have if you just coded it yourself.
jaggederest 1 days ago [-]
Those sneaky LLMs going 7 years into the past and committing as a human:
It’s obviously Claude 69 with time travel functionality, that’s too dangerous to release to public. They’re working on space-time limiting sandbox to prevent these issues.
six_seven 1 days ago [-]
Its all fun and games until the Claude-who-remains hunts you down
jaggederest 1 days ago [-]
Just remember kids, never immanentize the eschaton.
vegnus 1 days ago [-]
You're not reading it right. The bug was found using a vibecoded fuzzer.
12j3afAv 1 days ago [-]
I wonder from where Claude stole this fuzzer.
pjankiewicz 1 days ago [-]
Or it used something called an "analogy" which is a valid way to solve new problems.
I think they're talking about the misconception that LLMs can only ever regurgitate their training data verbatim enough to constitute mass copyright violation. And that that's therefore "stealing"
Rendered at 05:19:36 GMT+0000 (Coordinated Universal Time) with Vercel.
Edit: And there was discussion about this back in 2024 as well
- - -
From: Anthony Hurtado <[redacted since hn has no scrape protection]>
vpk_read_packet() divides vpk->last_block_size and (par->block_align - vpk->last_block_size) by par->ch_layout.nb_channels without checking for zero.
While vpk_read_header() validates nb_channels > 0, the codec parameters may become zero through format probing misidentification (VPK probe score is 2/3 of AVPROBE_SCORE_MAX) or codec parameter reset, causing SIGFPE.
Fix by:
- Checking nb_channels != 0 before division in vpk_read_packet
- Returning EOF for empty last blocks (last_block_size == 0)
- Validating block_count > 0 in vpk_read_header
- Validating last_block_size <= block_align in vpk_read_header
Found by fuzzing with libFuzzer + AddressSanitizer. Reproduces with 10 distinct inputs.
[patch redacted for brevity]
Maybe it uses other heuristics like detecting if someone lives in the AI agent world
I doubt the real bots have any trouble bypassing it.
My ungrounded speakers have a tendency to pop and make noise when they come out of sleep (sleep? on a speaker? fuck you logitec) and every time these "simple" CAPTCHAs come up, even if I pass without solving their logic puzzles, I hear the speakers activate as the Javascript on the page is figuring out what kind of audio setup I have by playing a silent sound file.
The default Anubis config isn't really a problem for any devices I've tried, but the FFMPEG Anubis setup is quite extreme. I seem to be served the extra-difficult Javascript challenge, as well as a high-difficulty challenge, that takes even powerful computers quite a long time to complete.
Could just be countermeasures to the hug of death every website gets when they get linked on HN, though, but someone would need to set up auto-scaling for that.
It's near instant on desktop (Windows/FF/7950X3D) and I wouldn't expect the delta to be that large against a modern mobile device.
Seems like it. It loaded near instantly as well from my Android smartphone using Firefox.
And contrary to the directions, refreshing the page DOES help.
They're paying for electricity and taking data without paying for it. It seems to me that they're paying for it exactly the same way everyone else in AI did.
This makes no sense to me as in my experience, you click the checkbox and then it verifies you without extra steps.
Having my CPU go up for a while is nearly frictionless on the other hand. Worst case I'm stuck in a loop and the site isn't loaded when I get back to it, which is better than being stuck in a captcha loop and then not getting to the site.
Of course, not having to do any of that would be even better. I wish the concept of ZeroNet had caught on, where everything is hosted and served peer-to-peer. This gives you basically zero hosting costs and you are immune to DDOS.
oh is that why my raspberry pi 5 can't browse websites anymore without freezing for a minute.
> how mild the reactions are to Anubis compared to the people reacting to seeing one singular Cloudflare captcha checkbox
cloudflare requires me to work for them
https://amber.agentic.church/web/lists.ffmpeg.org/2026082807...
Apparently, not the fonts they use for icons.
I get this crap when browsing on desktop a lot as well, principally because I stubbornly use Firefox as my main browser, and I habitually use a VPN when I connect my laptop to unsecured or even secured-but-accessible-to-large-numbers-of-people WiFi networks.
Like, seriously, bot detection "specialists", fuck off: I'm not a bot but your bot detection software IS shit, and I DO resent your shit software draining my battery and getting in my way. Learn to do your jobs properly, will you?
And don't come crying to me about how the problem you're trying to solve is "hard". I don't care: you chose it, you chose to considerably worsen the web browsing experience of millions of people globally, nobody made you. So go and find a different job if you're incapable of doing the one you have.
And if it's so "hard" why does your entire solution seem to be predicated on anyone's a bot if they're not running Chrome, or they are running an adblocker, or they appear to be from an unusual country that doesn't match their system language? Seriously, is this the level of sophistication you hacks operate at? To solve your "hard" problem?
You are extremely lame. Get out of my way.
Unfortunately, if you let all the bots in, they overwhelm your servers, and then nobody can access the website.
I had some git hosting up for a while, and was serving hundreds of qps and several terabytes per month. I can only imagine want significant sites are serving.
Such a non-argument.
I'm expecting people to create better, more effective, and less intrusive anti-bot measures. Measures that accurately detect bots but don't exclude real people from the web simply because of the browser they're using, or the country they either appear to be in or are in fact in, for example.
I've only seen 1 or 2 that know what they're doing. One's at lemmy.world ... just hovering over it is verified ...
Maybe now that ffmpeg is using Forgejo, the ball won't be dropped like this as often. Or there'll just be a five-digit number of open pull requests instead.
I understand why people choose it, but if I see the catgirl I close the tab - same is I get the test from cloudflare.
I was startled by the girl the first time as well, but once I learned what it is, I accepted it. Like a garden gnome on someone's front yard.
Isn't it easier to just .. start loving it instead?
You may want to check upthread how loved it is :)
I, for one, don't hate it, but I hate what it represents and see its existence (rather, the reason for its existence) as a defeat for the web.
(By the way: Jackal, not cat.)
The big problem preventing this approach from working for numbers is that it's just so cumbersome there. Most of this is because all the arithmetic operators are bundled into a single Num typeclass, and `fromInteger :: Num a => Integer -> a` has a type that's impossible for a "non-zero number" wrapper to satisfy.
Agda is the most mature dependently typed programming languae (having been around since the 90s – it is basically Haskell on steroids), but has a more proof-assistant flavor than an actual programming language flavor. Opus & Fable write Agda quite well, so LLMs can understand dependent types.
You need range proofs to be 100% safe, and then you can as well use the regular type because invalid values will not occur.
The only thing left for you to do should be directional decisions. The LLMs should pause and rope you in if the fix involves directional/invariant changes.
We wont stop using AI.
We will use AI to check AI.
Of course this is crazy, but it will also unlock pretty insane scaling and productivity and ultimately we will manage it on either end via requirements and tests.
Insane scaling of bloat, bugs, and technical debt I'd say.
> We will manage it on either end via requirements and tests
It is so crazy that this is being touted as a sane strategy. When I was a much worse programmer, I tried to write a big complicated string manipulation function to take two types of scripts in a language and add diacritics. I had the requirements very clear. I had the tests very clearly with all the edge cases. But I didn't have a good and clear picture of how to attack the problem which was quite novel for me. As I got closer to passing all the tests it got exponentially more unruly and confusing. And nearing the end I was frantically changing little bits here and there wincing and praying and hoping the tests would pass. "Please work! Come on!" Then when I got close enough, I could never ever think about touching that mess again.
I was a below average programmer then throwing myself at some novel problem I didn't understand. Throwing LLMs that produce below average code at novel problems and relying on tests and requirements is not where we want to go to make real progress.
(Years later after much learning and coding myself I was able to redo the function in a totally different way. This time I actually understood how to attack the strange problem and made something clean, clear, and robust that just worked. The tests then become a secondary guardrail, not the main force of correction.)
We are seeing such a massive regression from what we've learned over the years of CS.
Generating code automatically when you're not even quite sure what it is or even should be doing is insanity.
You just described every legacy codebase. Many of which are widely used and do a lot of sales. You dont need a clean codebase to have a valuable product.
>It is so crazy that this is being touted as a sane strategy.
Re-read what I said. I literally called it crazy.
It is the same dynamic that gave us customer service from some call center in India. Why would companies do this? Customer service got worse. Are they stupid? No, it's just worth it. The quality goes down but the business can scale more so it doesnt matter.
AI will absolutely be good enough at doing things that we'll happily accept some jankiness at times so that we can devote an extra 3000 hours per year per person to other things.
Im not even suggesting its a good thing. I just think the incentive structure dictates it. You're not going to have time to maintain a small slice of some service by hand.
From that standpoint, it's not a crazy setup security-wise. Maybe still crazy for development.
I had Fable add a new subcommand to our internal CLI tool. I reviewed and tested it locally and had to suggest several fixes that I feel like I wouldn't have had to tell a human senior engineer to do. When it finally submitted the PR, I had it on a loop waiting a few minutes for comments on the PR, then assessing/addressing/replying-to/resolving them, and then repeating again until all AI reviewers were okay with it. It ended up going through dozens of revisions and ended up with 160 comments left on the PR.
Are you implying that all code writing by LLMs atm is bug-free?
But, to the underlying question, obviously as we automate more and more of our work, of course we provide less and less value. We're heading towards a future where selling thought for money isn't going to work so well.
It used to be considered a quality of good code that there would be less code, not more.
Some people always tryin to get the highscore on golf.
We didn't stop using syntactic programming languages we used code to check code.
Not sure it's really crazy at all. It's been an abstraction for programmers probably since we stopped soldering transistors to each other.
But if you don’t actually read it…
They find bugs but whether they save time is nowhere near as clear as you try to insinuate here.
Generate multiple solutions- they do not to work 100% correctly. And than I check which I would prefer. Which is more to our applications taste.
And than I would take the vibe output as a kind of a ‚plan‘ which I use to implement but not follow 100% and at the end I take my solution and review it. I gain speed with that because I often can quickly see the pros and cons of a solution way better than when I would manually do it and hang on a major roadblock and also I even see such roadblocks in the vibe output - it’s mostly the part with an unnecessary amount of new code that looks nonsensical.
No big deal? It’s not like it’s free… tokens cost money.
But yeah, I guess it can be used to increase certain aspects of quality by letting them go wild. But I think I mostly hear about security or crash issues. In my experience they don't outweigh the number of other issues they cause. Like UI bugs. I've seen more than one service constantly rolling out features that are completely broken, just to have a completely new, still broken, solution available the next day.
There are also non security bugs that don't have exploits but just make the user experience worse.
A.I. could also be used to port C/C++ codebases to Rust, which isn't economically feasible at the moment.
it may not be an issue for ffmpeg, but it might be for an application that bundles ffmpeg.
this isnt reddit
Two months and 1.1k+ commits to rediscover an already known bug IS funny. That doesn't mean I hate the person or want to discredit their work. You're reading way too much into a comment :)
I guess sometimes developers think they "know" some variable won't be zero, but unless it checked explicitly or by the compiler, that shouldn't be trusted.
If you’re accepting large false positives rates: yes.
If you want users to take your warnings serious: no.
(Nitpick: you certainly don’t want to flag _all_ of them. Divisions by non-zero constants definitely should be excluded, for example (integer division by -1 can lead to overflow, but that would be a different warning))
Division can 'go wrong' for certain inputs, but it's not just division. In C, signed integer addition, subtraction, and multiplication, all give undefined behaviour on overflow.
As 'Someone' already pointed out, it's not helpful to just flag all uses of the division operator, or of other potentially dangerous operators. Minimising false positives is one of the core challenges of program analysis.
But this bug feels like something an LLM would flag as a major finding but turns out to be completely benign.
Update: I tried to look into the fuzzer but it is hard to get past the AI blabb. Can someone please explain to me what it does beside being structure aware?
The implications of the crash depends entirely on the implementation of the process it crashes. If I use ffmpeg as a library it'll crash my process upon processing the offending file. How is my process designed? How is every process that uses ffmpeg designed? You don't know, therefore you can't say that it's a "self-dos" in every case even if you know that it is in some cases.
Maybe I am clever enough to have read up on the history of ffmpeg vulnerabilities before deployment to an attacker-facing service and have designed a solution where a crash has minimal implications, but maybe I'm not, and haven't. It's beside the point.
It is not enabled by default. It is used only in video games, which input files are fixed set of asset that came with the game.
It can be a crash, yes. but the typical user of this codec won't care.
My current opinion on LLMs is that they are superhuman in that they lack fatigue, they have close to full knowledge across all subjects which are known to humans at least publicly, and the fact that you can vibe code a harness to look for bugs in a famously complicated C codebase is intern level stuff and hardly news.
Smart aspiring blackhats will be targeting tmux next, both with light llm jailbreaks, light supply chain attacks (web search results) and LPEs within certain environments which weren't particularly useful before but with agents running on auto mode for hours become a very valuable springboard. I'm not sure on the quality of tmux code but I know its written in C and is very complex and was not at all designed to defend against this type of threat.
So you’re back to depending on the agent to execute code locally. at which point you’ve already compromised the system so don’t need a tmux bug.
I’ve spent a lot of time in tmux. Including writing a frontend for it. So I’m probably more familiar than most. And I hear a lot of people say tmux (specifically) is a vulnerability because it’s written in C. But I struggle to see how it’s any more of a vulnerability than (for example) coreutils. Or any other piece of software for that matter.
(There actually was one privilege escalation bug in tmux, but it actually seems like a distro packaging error. The distro setgid the executable so the resulting shell inherited the additional group. This didn’t require any exploit, that’s just how child process inheritance works.)
persisted malicious code running in your tmux process that you don't know about is probably not where you want to be, for obvious reasons.
If you wanted to booby trap a repository then you’re far better off with a prompt injection attack.
It seems like this would have been pure fantasy not that long ago though. So why isn’t it noteworthy again? I don’t really follow what you’re complaining about.
https://ffmpeg.org/pipermail/ffmpeg-devel/2024-November/3355...
It’s surprisingly simple to setup, and the hardware requirements are pretty small for a private or small forge, as it’s usually a relatively small number of users/repos/etc.
Git is a DVCS. I know many people only ever used Git through Github and forgot what the 'D' in DVCS means but whether or not they remember what the 'D' stands for, running your own Git server is trivial. Especially in this day and age of LLMs were you can just ask: "Clone this repo and convert it to base Git repo and serve it on the LAN PLZ KTHX".
The result is going to be more stable than Github and, arguably, more secure too.
maybe we'll just see them remove support for these long-tail formats the way linux has been removing drivers for similar reasons https://www.phoronix.com/news/Linux-Retiring-Moxa-Driver
Next week, I'll apply for the cyber and I suspect I'll find something similar.
Right now, it's just annoying and thanks the OpenAI cyber was much easier to get access to.
I understand the utility of this though, I just don't see this particular bug and something that would be particularly difficult o find pre LLM era.
The fuzzer found the bug before any humans did, so there is a mismatch of developers who could find this bug and those who did (without an LLM-coded fuzzer).
The value of the fuzzer continues long after it found this one bug.
It's worth nothing that in the bug discussion thread, the bug fix author pointed out that it's not easy to set up the config then call the functions in the right order. Your comment assumes that the reader has enough context to read the code and build the finite state automata in their mind. The bug fix reporter's comments suggest that you are assuming things which you shouldn't assume.
The bug I located in SystemD was literally a null ptr exception. All they had to do was perform a null check on a cstring but they hadn't.
I don't see the utility of reporting an LLM made fuzzer finding bugs that could be found by a lint rule or static analysis. I will appreciate a post about an LLM fuzzing software to find unique corner cases, which I predict will happen soon, in ACL controlled systems caused by policy shadowing.
There's no viable way to statically prove at compile-time that these variables will never become zero at runtime, ultimately forcing a system of endless runtime checks (be it software or hardware)... which is why processors already throw exception interrupts when division by zero is attempted.
An alternative https://en.wikipedia.org/wiki/Projectively_extended_real_lin...
The projectively extended real line defines division by zero, no reason you couldn't have a floating point type that implemented it.
>There's no viable way to statically prove at compile-time that these variables will never become zero at runtime
strongly typed programming languages like Ada allow for types which have ranges such as disallowing zero -- but also any arbitrary thing like you can create a floating point "degrees" type which is [0.0, 360.0] or any other ranged type
They only suggested a basic guard, chich can be useless if this case never happens
Generating correct input to get deep into the call stack and then finding something is the hard part.
https://code.ffmpeg.org/FFmpeg/FFmpeg/commit/8eda3c7f91e1a5b...
> This is a bug found with our fuzzer: https://github.com/daedalus/fuzzer/