Command Code is live now. Try the first coding agent with taste.
04/07/2026
18 min read
Can an AI Pentest Replace Human Pentesters?

Every second an agent waits for infrastructure is a second it isn't doing useful work. When we started Codexa, a typical container took eight to twelve seconds to become ready. Today the median sandbox is ready in 180 milliseconds. Here's how we got there.
Lazy-loading container images
Most images are huge, but agents only touch a small fraction of the files inside them on startup. Instead of pulling an entire image before booting, our runtime mounts it lazily and fetches file blocks on demand. Frequently used blocks are cached on every host, so popular base images are effectively free to start.
A scheduler that thinks ahead
Our scheduler keeps a small pool of warm, pre-initialised sandboxes in every region and predicts demand from recent traffic. When a request arrives, it's usually matched to a sandbox that's already waiting, with only your code and secrets injected at the last moment.
Snapshot restore instead of reboot
For stateful workloads we restore a memory snapshot rather than booting from scratch. Restoring a 2 GB snapshot takes less time than starting a fresh Python interpreter and importing a few machine learning libraries.
The results
Median cold start: 180 ms (down from 9.4 s).
p99 cold start: 640 ms across all regions.
40% lower compute cost for bursty agent workloads.
We'll keep pushing these numbers down. If you enjoy problems like this, we're hiring infrastructure engineers in New York, Stockholm and remote.
Table of contents
Key takeaways
What is manual penetration testing?
What is AI pentesting?
AI vs. manual pentesting example
Authors

Lauren Volpi
Marketing
Share this article



