The "reviewing product output instead of code" framing is the right direction for where development is heading, and the per thread VM model makes sense for isolation.
One dimension worth thinking through as you scale: the security surface of cloud-hosted agents is meaningfully different from local agents. Local agents (Cursor, Claude Code) have access to your local filesystem and credentials. Cloud agents have access to your cloud credentials, your production-adjacent infrastructure, and potentially your CI/CD pipeline and they run without the developer watching.
The MCP server porting during onboarding is the piece I'd think hardest about. MCP servers can have write access to config files, and the STDIO transport has a documented unsanitized parameter passthrough vulnerability (AVE-2026-00060, corroborated by OX Security and Microsoft) that affects Python, TS, Java, and Rust SDKs. When you're running hundreds of agents concurrently in the cloud, a single compromised MCP server has a much larger blast radius than a local one.
The "what the agent writes" security layer is separate from the "how the agent runs" security layer. Hoplite solves the second. Scanning what the agent wrote (SAST, secrets detection, dependency audit) before it gets merged is the complementary first layer. SafeWeave runs as an MCP server inside the agent's environment for exactly this works locally and in cloud agent setups.
Congrats on the launch. What's your current approach to credential scoping for agents that need cloud access?
scottydelta 1 days ago [-]
Trying to wrap my head around how it differs from my current on-the-go setup that is claude code. On claude's phone app or web app I can choose a repo, ask it for a feature and it writes the code, runs my tests + add more tests and creates a new branch. Then I can click on create PR or configure claude code to auto create PR.
I also have another setup which is a self-hosted docker compose behind my vpn with one container with claude code agents managed using agent of empires[0] and another container with playwright with sse. Using this setup, my agents get access to actual browser where it can test things live and I can access the app started by agent on domain:port. This is something I don't get with Claude.
[0] https://github.com/agent-of-empires/agent-of-empires
BenceRed 1 days ago [-]
That setup is pretty much what we're trying to offer with Hoplite!
Using us means losing freedom and control with regards to infrastructure, however we think that's a tradeoff people would want to make in exchange for easier onboarding and a more polished experience.
scottydelta 1 days ago [-]
Makes sense. It took me some experimentations with a few open source libraries and docker-compose to arrive at my setup. My setup also requires me to access agents via a terminal app on my phone so it's nice to have a web app like your offering.
Are you offering browser access to the agents in your setup?
BenceRed 1 days ago [-]
Yes, agents have a persistent Chromium session they use via the agent-browser CLI. Typical workflow would involve starting the preview, seeding data, then the agent going through the old and new UX flows for a before + after view. We've also got some optimisations around saving the aforementioned flows in a QA library, so that they can be replayed without needing an agent to run through it all again.
sebmellen 13 hours ago [-]
At the risk of replicating the classic Dropbox post (why would you need this when you could just use rsync?)…
I have a dev box with 96 GB of RAM, 2x4 TB NVMe drives, and an unbelievably beefy AMD CPU. This box costs me less than $150 per month and is so hyperlocal that I can log into it and use it as a remote desktop, while also using it as an always-on server that I can use to run T3 code and tmux and so forth. I can then connect to it from my laptop or my phone using Tailscale and prompt using the T3 Code or Remux mobile apps. Voilà — I have my own outsourced development center.
In this setup, my agent can handle everything: previews with a NixOS environment, unlimited threading, “autofixing” (which is just a loop between my agent and Copilot review comments), etc.
But it requires a LOT of custom setup/tooling so that my local environment works with my agent.
Why am I telling you this? Well, I've tried a number of serverless or ephemeral VM-type solutions, and it turns out that once you're working on "real code," you can't use ephemeral micro VMs reliably because your code starts interacting with too many different dependent services. You have to run migrations, so that your tests run properly, and to do that, you need to pull five different Docker images, and it goes on indefinitely. Eventually, the overhead of making little micro VMs is so high that it makes much more sense to take a monolithic approach to development and have a persistent workstation. You can still use things like worktrees, which allow you to massively parallelize your work, but you're building off of a shared local drive and cache.
So I believe there's a place for something like Hoplite with simpler software, but the problem is that the minute you get beyond toy software, it becomes really hard to test, scale, and deploy everything in micro VMs. There are also other companies that have tried this approach (like https://shipyard.build, although I think they had a slightly different philosophy from what you're doing) and I don't know that they've been massively successful.
What is it that you're doing differently that will allow Hoplite to succeed? How do you think that you'll compete against the legacy players in this space and the more full-spectrum players like Devin, et al.?
BenceRed 11 hours ago [-]
Re: the dev box, it works very well for individuals and small sized teams, but starts to become an operational burden past a certain size. Our ideal customer is one who has a ton of engineers and wants great multiplayer/observability, as the case for Hoplite becomes a lot clearer -- "Run through project setup once, then onboard all engineers with one email (and they can bring their entire local setup with one CLI command)".
On the point of microVMs, agreed that it's a very difficult problem to solve. Luckily sandbox providers are continuously improving their APIs to make this slightly easier, but I wouldn't be surprised if we need to migrate over to AWS Lambda MicroVMs and roll a lot of the orchestration logic ourselves. Our goal is to get our P95 project setup time (i.e. connect -> fully running in the sandbox) to around 5 minutes, most of which we imagine being dependency installation. This is one clear point of differentiation where if we nail it, we'd be leagues above the rest of the competition.
The legacy players such as Devin, Cursor, and Factory are certainly well entrenched in their market position, but this space has the unique advantage of completely reworking how it operates every 6 months. These existing tools need to balance keeping up with new user demand for features, while also maintaining the old legacy workflows for their existing customers. We're lucky in that we can now build a product that we believe resembles how the majority of development will work ~1 year from now, meaning we have a lot more flexibility in how we can move forward.
And fundamentally, outside of large enterprise features like on-prem deployment, the key differentiators are 1) UX, 2) cost, and 3) harness performance. We can certainly win at the first, are at parity with the second, and likely struggle at the third (need to do benchmarks/evals -- if those go poorly then we'll transition from the custom harness to using the first party Claude Code/Codex. So quite fixable.)
igorcardines 8 hours ago [-]
[dead]
pelagicAustral 11 hours ago [-]
The pricing seems draconian... why would anybody choose to pay this way when you can pretty much do the same with exe.dev?
tiborsaas 10 hours ago [-]
Github already does this, there's an agents tab.
kristianc 1 days ago [-]
>> It opens a pull request
>> Then keeps iterating as review comments land, in the same thread, with the same context.
My experience, particularly with Sol is that agents are generally really bad at knowing when to stop 'iterating' and will continue covering off ever-more obscure edge cases. How does Hoplite solve for this?
BenceRed 1 days ago [-]
We've observed this pattern as well, and have counteracted it by keeping the tasks very finely scoped. For example, "The test(api) check has failed. Fix it, then immediately commit and push." -- or -- "The following comments have been added by reviewers. Resolve each one, then immediately commit and push."
This specificity helps Sol stay on track (most of the time). It doesn't work as well when the comment questions a complex piece of the architecture though.
fishtoaster 1 days ago [-]
Took me a minute to see the value - my first thought was "this is just cursor's cloud agents..."
But the key thing here for me is "Every sandbox boots your app on a live URL." Cursor doesn't easily have that, and that's what would allow me to ditch my local env entirely - the ability to actually try out a PR without needing to check it out locally.
So on that note: how does that work? We've had trouble with getting our dev env running in other cloud envs because it requires a few things (clickhouse, localstack, pg, etc) running which we manage via docker compose locally.
Also, some minor pricing feedback: it'd be really great if there were a version with pay-as-you-go and a cheaper fixed cost. I think your 99/seat/mo model is fine for professional work, but it's a lot to commit to for personal work.
sergeyk 5 hours ago [-]
Check out https://superconductor.com: supports docker and env setup is done for you by an agent when you connect repo
BenceRed 1 days ago [-]
Modal just released some features that would allow users to bring custom Docker images, and I'm working on getting your exact use case supported! Aiming to get it out by the end of the week.
Noted the pricing feedback! We're still figuring out exactly what works best so it's still very much so subject to change.
r5Khe 1 days ago [-]
Looks neat! I've been using Amp (https://ampcode.com/) for a while (Which seems to be doing something very similar), and I really appreciate this type of workflow. One thread = one VM feels like a solid model going forward.
yoanwaidev 23 hours ago [-]
Same instinct locally with git worktrees: one agent per worktree so they do not thrash the same checkout. Cloud VMs buy isolation and scale; worktrees buy cheap isolation when you already have a machine. I end up using both patterns depending on whether the bottleneck is compute or just not stepping on each other.
BenceRed 1 days ago [-]
Agreed. Per-thread VMs are quite similar to how local agents use worktrees to avoid cross contamination, but with the added benefit of being able to easily scale up/down compute requirements on demand.
mellosouls 1 days ago [-]
Firstly: good luck!
I've been wondering what the alternatives to things like Github Copilot Cloud and Codex Cloud might be, especially ones that might be flexible wrt models, and this seems at least to have some of those behaviours.
If that perception is correct, please would you explain what it offers against those sorts of services (those in particular) and how the pricing compares - eg. their base levels are $20 a month, yours starts at a higher level - I can see there seems to be more brought in from the local IDE world (and similar), which seems very useful compared to the standard "prompt against repo, repeat" of the normal cloud agents but it would be useful to understand the targets and intents.
BenceRed 1 days ago [-]
Thank you! Regarding models, as you said we're not locked into a specific provider, and are able to offer open weight models like Kimi K3 and GLM 5.2
Our pricing is higher than other providers because we do not upcharge on token or sandbox costs. We believe that people should be running as many agents as they possibly can handle, and an upcharge would create a monetary incentive for us to say that, when it's a genuine belief we hold.
We also offer features out of the box that would usually be behind enterprise gating (e.g. sandbox baking).
mohammedmsgm 12 hours ago [-]
The custom harness bet is the interesting one
not using Codex/Claude Code trades a lot of free improvement for control you may not need yet, so I'd watch whether that pays off before the underlying models plateau
BenceRed 12 hours ago [-]
We're going to be investing pretty heavily in evals/benchmarks over the next couple of weeks, so that should give us a much better understanding of how our custom harness stacks up to the official ones.
xander_north 1 days ago [-]
Really interesting. I've been working on building out more self-sufficient agents with harnesses to match, and this feels like the next logical step. I'd love to try this.
Curious how you guys are thinking about the difference between project-specific and project-agnostic harnesses and tooling. For me, it feels like a lot of the work is project specific and I'm not sure how to abstract that.
P.S. The code is not working for me.
BenceRed 1 days ago [-]
In general, we've found that over the past couple of years agent harnesses have gotten much less restrictive, allowing the agent freedom to choose its own way of doing things. It seems like the project agnostic vs. specific harnesses will follow that trend. The pattern will work well for the current gen models, but eventually Fable 7 will be able to intuit how it should approach a specific project very well, at which point the challenge is making sure it has the tools to do so.
And I've made some changes to the coupon, does it work now?
xander_north 24 hours ago [-]
It does, thanks! Looking forward to trying this out.
ChrisMarshallNY 23 hours ago [-]
Good name, with The Odyssey out, and all. I guess that you could think of the agents as "footsoldiers," and whatnot.
Good luck!
BenceRed 22 hours ago [-]
We use 'phalanx' internally!
Bnjoroge 1 days ago [-]
What’s the experience like going from an active on-going thread to a cloud-hosted one? I dont wanna always work on the cloud, and want my current setup to be exactly the same as in the cloud, and should be pretty seamless. Only folks i’ve seen solve these are folks who run the sandboxes locally and take that to the cloud like smolvm/microsandbox.
BenceRed 1 days ago [-]
This is still something we're working on making seamless. The current approach is to install our MCP server and ask the local agent to start up a new thread on Hoplite when you want to transition to the cloud, but it doesn't carry over file system changes. (Unless you first push the contents to a remote branch, at which point the Hoplite agent can pull it down.)
Bnjoroge 1 days ago [-]
gotcha. yea ideally you take a full live snapshot of my current state, untracked files, processes etc, and resume them in the cloud. Congrats on the launch! A mobile app would also be nice
lionls 1 days ago [-]
Amazing work, I am currently in the process of building something similar on my server for personal use, but yours looks really promising. Especially running sandboxes on your own can be tedious. Why have you opted for Modal instead of Firecracker or a similar micro VM solution?
Best of luck to you!
BenceRed 1 days ago [-]
Modal has a lot of small niceties that made them easy to implement, such as filesystem snapshots and programmatic build images. But I did see that AWS recently launched Lambda MicroVMs, and since we're an AWS house we may transition to using them.
My main issue with Modal is that their autoscaling is not as good as Daytona's. You have to stop the machine, resize, then start it, which takes ~3s and terminates all running processes. Daytona supports scaling up (but not down) without stopping the VM.
Also would recommend checking out ColeMurray/background-agents if you're planning to self host. Very good alternative! And the team behind it are great
lionls 1 days ago [-]
Thanks Bence, for your thoughts there. Will definitely take a look.
asdev 1 days ago [-]
just a data point, at my company we are building this internally. if you're targeting people building from scratch it might work, but there's no way you can port any somewhat mature infra stack, nor will the org want to. you'll need to deal with the variable complexity of everyone's dev environment which already doesn't work locally for thousands of different reasons.
BenceRed 1 days ago [-]
Agreed that at the moment it's a very difficult problem, but one we're looking to solve! I think it becomes a no-brainer for most people if we're able to give each agent a replica of their production stack.
What does your current setup look like? And are you using an open source solution like OpenInspect for your in-house version, or building it from the ground up?
asdev 1 days ago [-]
Building from ground up using OpenAI Agents SDK. We already have custom in house Cloud Development Environment, so the effort is just to "agent-ize" though which is not small. That's why I feel any team with any sort of infra support likely won't buy your product, since they already have the tribal knowledge to set this up. Newer teams might. But overall porting people's dev envs into the cloud is a tarpit problem(IMO), I was interested in this space too but decided against it for that reason. Happy to be proven wrong though!
BenceRed 1 days ago [-]
I think for use cases like that, we'd offer on-prem deployments (similar to Factory), potentially coupled with a FDE. Still need to do a lot more research into the enterprise space.
FailMore 1 days ago [-]
Looks interesting. Does this mean I would use this as my day to day harness? Or is it something additional to an established workflow?
BenceRed 1 days ago [-]
You can do either. If you don't want to migrate over fully, I'd recommend setting up an automation to fix Sentry/PostHog issues as they come in. You can get a good feel for the platform and how it fits into your workflows that way.
We also have an MCP server that you can use to delegate tasks (e.g. research, debugging, SRE work) to Hoplite via your existing local setup.
mkagenius 1 days ago [-]
If you ever need to switch sandboxes, would be happy to chat.
BenceRed 1 days ago [-]
At the moment we're using Daytona as a redundant fallback in case Modal experiences an outage, but they have very stringent limits on how many resources we can consume concurrently. We're evaluating adding a second provider to help ease this so would love a chat! Feel free to email bence [at] hoplite.sh
This one: https://www.daytona.io. Their platform was OSS for a long time but they decided to go closed source recently.
abtinf 1 days ago [-]
Why would I use this over exe.dev?
BenceRed 1 days ago [-]
exe.dev works quite well for giving an agent a computer and managing it remotely, but seems to require a fair bit more configuration to achieve parity with what we offer out of the box. Namely automations, PR autofix, visual QA, and general UI/UX polish.
I think it comes down to whether configuration or ease of use is valued more, and Hoplite favours the latter a bit more. (They shouldn't really be mutually exclusive, but we have a long way to go before we're happy claiming that we match/beat self-hosting in that area)
docheinestages 1 days ago [-]
Suggestion: showing an actual screenshot or video of your app is a much better indicator of effort than a generic Claude made animation. I've seen AI slop landing pages on far too many YC-backed startups. Not saying yours is one, but parts of it smell.
BenceRed 1 days ago [-]
Agreed 100%. We've been working with a designer on a complete redesign of our landing page to avoid that vibey-smell.
Agreed, we're working with a designer and are going to be fixing this very soon. Our main focus has been on making sure the product itself looks and feels very good to use -- probably not the best approach from a marketing POV.
Almost every possible name that is not a portmanteau, made-up word, or combination of N words has a naming "conflict". This is not an interesting thing to say and I wish that people would stop saying this on everything anyone ever posts on HN. It almost feels like it should be against HN rules to point out that something else shares the same name with no additional statements.
LoganDark 1 days ago [-]
For one, it's interesting to me because I've already known Hoplite for years as nothing to do with AI. For two, I'm not sure how sharing that is so egregious it should be against the rules? Is there an interpretation I'm missing of my original comment? Does pointing out another Hoplite get interpreted as disparaging or accusatory in some way? Does it go against intellectual curiosity?
rytill 1 days ago [-]
The term "conflict" implies "there is a problem here". I would say it's very lightly disparaging, because it implies the author didn't even do basic research on other things that are also named the thing they decided to call it.
I actually checked HN rules and just saw this:
> Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting.
So, there you go, I guess.
e55h5reh 22 hours ago [-]
you are complaining about yourself
BenceRed 1 days ago [-]
Yeah, luckily they're in quite a different domain to us -- hopefully shouldn't have too much trouble winning the SEO battle
LoganDark 1 days ago [-]
I don't imagine they get too many Google searches anyway, but they are incredibly popular in competitive Minecraft spheres.
rirze 1 days ago [-]
Good luck, `hoplite` triggers a ton of Minecraft media for me and that name has been around 10 years.
Ozzie-D 19 hours ago [-]
The bet that devs shift from reviewing code to reviewing product output is the most interesting part of this. Most teams haven't really internalised that yet — they're still doing line-by-line code review on agent-generated PRs, which doesn't scale and arguably misses the point. If the output works correctly and the tests pass, the code style of an agent matters a lot less than people think.
The hard part is defining what "works correctly" means in a way that's automatable. Visual diffing helps but it's brittle for anything beyond static layouts. Curious how you're handling cases where the correct behaviour is contextual rather than pixel-perfect.
ThomasSchijf 14 hours ago [-]
I use syns.dev. The agents share one plan up front, so there is less to reconcile at review.
sebmellen 13 hours ago [-]
Yikes. Hello, Claudebot.
abratabia 1 days ago [-]
[dead]
Rendered at 21:01:39 GMT+0000 (Coordinated Universal Time) with Vercel.
One dimension worth thinking through as you scale: the security surface of cloud-hosted agents is meaningfully different from local agents. Local agents (Cursor, Claude Code) have access to your local filesystem and credentials. Cloud agents have access to your cloud credentials, your production-adjacent infrastructure, and potentially your CI/CD pipeline and they run without the developer watching.
The MCP server porting during onboarding is the piece I'd think hardest about. MCP servers can have write access to config files, and the STDIO transport has a documented unsanitized parameter passthrough vulnerability (AVE-2026-00060, corroborated by OX Security and Microsoft) that affects Python, TS, Java, and Rust SDKs. When you're running hundreds of agents concurrently in the cloud, a single compromised MCP server has a much larger blast radius than a local one.
The "what the agent writes" security layer is separate from the "how the agent runs" security layer. Hoplite solves the second. Scanning what the agent wrote (SAST, secrets detection, dependency audit) before it gets merged is the complementary first layer. SafeWeave runs as an MCP server inside the agent's environment for exactly this works locally and in cloud agent setups.
Congrats on the launch. What's your current approach to credential scoping for agents that need cloud access?
I also have another setup which is a self-hosted docker compose behind my vpn with one container with claude code agents managed using agent of empires[0] and another container with playwright with sse. Using this setup, my agents get access to actual browser where it can test things live and I can access the app started by agent on domain:port. This is something I don't get with Claude. [0] https://github.com/agent-of-empires/agent-of-empires
Using us means losing freedom and control with regards to infrastructure, however we think that's a tradeoff people would want to make in exchange for easier onboarding and a more polished experience.
Are you offering browser access to the agents in your setup?
I have a dev box with 96 GB of RAM, 2x4 TB NVMe drives, and an unbelievably beefy AMD CPU. This box costs me less than $150 per month and is so hyperlocal that I can log into it and use it as a remote desktop, while also using it as an always-on server that I can use to run T3 code and tmux and so forth. I can then connect to it from my laptop or my phone using Tailscale and prompt using the T3 Code or Remux mobile apps. Voilà — I have my own outsourced development center.
In this setup, my agent can handle everything: previews with a NixOS environment, unlimited threading, “autofixing” (which is just a loop between my agent and Copilot review comments), etc.
But it requires a LOT of custom setup/tooling so that my local environment works with my agent.
Why am I telling you this? Well, I've tried a number of serverless or ephemeral VM-type solutions, and it turns out that once you're working on "real code," you can't use ephemeral micro VMs reliably because your code starts interacting with too many different dependent services. You have to run migrations, so that your tests run properly, and to do that, you need to pull five different Docker images, and it goes on indefinitely. Eventually, the overhead of making little micro VMs is so high that it makes much more sense to take a monolithic approach to development and have a persistent workstation. You can still use things like worktrees, which allow you to massively parallelize your work, but you're building off of a shared local drive and cache.
So I believe there's a place for something like Hoplite with simpler software, but the problem is that the minute you get beyond toy software, it becomes really hard to test, scale, and deploy everything in micro VMs. There are also other companies that have tried this approach (like https://shipyard.build, although I think they had a slightly different philosophy from what you're doing) and I don't know that they've been massively successful.
What is it that you're doing differently that will allow Hoplite to succeed? How do you think that you'll compete against the legacy players in this space and the more full-spectrum players like Devin, et al.?
On the point of microVMs, agreed that it's a very difficult problem to solve. Luckily sandbox providers are continuously improving their APIs to make this slightly easier, but I wouldn't be surprised if we need to migrate over to AWS Lambda MicroVMs and roll a lot of the orchestration logic ourselves. Our goal is to get our P95 project setup time (i.e. connect -> fully running in the sandbox) to around 5 minutes, most of which we imagine being dependency installation. This is one clear point of differentiation where if we nail it, we'd be leagues above the rest of the competition.
The legacy players such as Devin, Cursor, and Factory are certainly well entrenched in their market position, but this space has the unique advantage of completely reworking how it operates every 6 months. These existing tools need to balance keeping up with new user demand for features, while also maintaining the old legacy workflows for their existing customers. We're lucky in that we can now build a product that we believe resembles how the majority of development will work ~1 year from now, meaning we have a lot more flexibility in how we can move forward.
And fundamentally, outside of large enterprise features like on-prem deployment, the key differentiators are 1) UX, 2) cost, and 3) harness performance. We can certainly win at the first, are at parity with the second, and likely struggle at the third (need to do benchmarks/evals -- if those go poorly then we'll transition from the custom harness to using the first party Claude Code/Codex. So quite fixable.)
My experience, particularly with Sol is that agents are generally really bad at knowing when to stop 'iterating' and will continue covering off ever-more obscure edge cases. How does Hoplite solve for this?
This specificity helps Sol stay on track (most of the time). It doesn't work as well when the comment questions a complex piece of the architecture though.
But the key thing here for me is "Every sandbox boots your app on a live URL." Cursor doesn't easily have that, and that's what would allow me to ditch my local env entirely - the ability to actually try out a PR without needing to check it out locally.
So on that note: how does that work? We've had trouble with getting our dev env running in other cloud envs because it requires a few things (clickhouse, localstack, pg, etc) running which we manage via docker compose locally.
Also, some minor pricing feedback: it'd be really great if there were a version with pay-as-you-go and a cheaper fixed cost. I think your 99/seat/mo model is fine for professional work, but it's a lot to commit to for personal work.
Noted the pricing feedback! We're still figuring out exactly what works best so it's still very much so subject to change.
I've been wondering what the alternatives to things like Github Copilot Cloud and Codex Cloud might be, especially ones that might be flexible wrt models, and this seems at least to have some of those behaviours.
If that perception is correct, please would you explain what it offers against those sorts of services (those in particular) and how the pricing compares - eg. their base levels are $20 a month, yours starts at a higher level - I can see there seems to be more brought in from the local IDE world (and similar), which seems very useful compared to the standard "prompt against repo, repeat" of the normal cloud agents but it would be useful to understand the targets and intents.
Our pricing is higher than other providers because we do not upcharge on token or sandbox costs. We believe that people should be running as many agents as they possibly can handle, and an upcharge would create a monetary incentive for us to say that, when it's a genuine belief we hold.
We also offer features out of the box that would usually be behind enterprise gating (e.g. sandbox baking).
not using Codex/Claude Code trades a lot of free improvement for control you may not need yet, so I'd watch whether that pays off before the underlying models plateau
Curious how you guys are thinking about the difference between project-specific and project-agnostic harnesses and tooling. For me, it feels like a lot of the work is project specific and I'm not sure how to abstract that.
P.S. The code is not working for me.
And I've made some changes to the coupon, does it work now?
Good luck!
Best of luck to you!
My main issue with Modal is that their autoscaling is not as good as Daytona's. You have to stop the machine, resize, then start it, which takes ~3s and terminates all running processes. Daytona supports scaling up (but not down) without stopping the VM.
Also would recommend checking out ColeMurray/background-agents if you're planning to self host. Very good alternative! And the team behind it are great
What does your current setup look like? And are you using an open source solution like OpenInspect for your in-house version, or building it from the ground up?
We also have an MCP server that you can use to delegate tasks (e.g. research, debugging, SRE work) to Hoplite via your existing local setup.
As the repo says no longer maintained
I think it comes down to whether configuration or ease of use is valued more, and Hoplite favours the latter a bit more. (They shouldn't really be mutually exclusive, but we have a long way to go before we're happy claiming that we match/beat self-hosting in that area)
I actually checked HN rules and just saw this:
> Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting.
So, there you go, I guess.
The hard part is defining what "works correctly" means in a way that's automatable. Visual diffing helps but it's brittle for anything beyond static layouts. Curious how you're handling cases where the correct behaviour is contextual rather than pixel-perfect.