Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
Published
The script
Panel 1
Alex: Hey team! I've got a project I think we should tackle. I'm benchmarking our AI models on real-world, enterprise codebases. It's going to be a game-changer!
Sam: Uh, Alex. Did you check if your benchmarks are ethical and legal?
Robin: You mean like the time you tried to use Houdini's lunch money?
Alex: Well... yeah, and that was a bit of a disaster, but still, it's a start!
Panel 2
Alex furrows his brow, looking at a codebase filled with technical documentation and files. He looks up, eyes wide.
Alex: Uh oh... This codebase is bigger than I thought. And don’t even get me started on the spaghetti algorithm.
Sam: Remember Alex, we need to be careful not to break anything. And we definitely need a backup plan.
Robin: Like a backup plan for when you get hacked by a piece of code?
Alex: Maybe... I’ll work on that too.
Panel 3
Alex’s laptop hums, the screen showing a mix of code and diagrams. He looks excitedly at the screen, a smile spreading across his face.
Alex: Alright, I’ve got this. I’ll start with the most stable parts of the codebase. Maybe we can use this as a way to train our AI to better understand enterprise code.
Sam: You know what they say, Alex. The best way to understand enterprise code is to write enterprise code. But be careful, the AI might start writing its own code in a way we don’t like.
Robin: And maybe we’ll need to have a meeting with the lawyers to ensure we’re not getting into any trouble.
Alex: I’ll get started right now. Let’s see how this goes.