Many companies live with internal applications they no longer dare to touch. The code runs but relies on an obsolete technology that is no longer maintained. There are no longer any security updates, and therefore vulnerabilities accumulate (passwords are stored with unsalted MD5, for example). If the question of a technical update has regularly arisen, the step has never been taken and the debt has accumulated. The company's codebase is therefore kept without maintenance, for lack of a better option.
However, on 16 July 2026, the American company Anthropic
announced having entirely rewritten
Bun en Rust using its Claude language model. In the era of generative AI, the balance has shifted. For multi-agent systems and the latest language models, the scale of the codebase has been reached. The
migration of a backend of 100,000 lines of code by AI models is a task
achievable in
two weeks.
The objective of the migration is multi-faceted: we update the technical stack, ensure that the migration only changes the syntax (and not the behaviour) of the code, and detect bugs or vulnerabilities along the way. The migration proceeds as follows:
We map the endpoints of the legacy code: these are the entry points from which the migration operates
For each of the endpoints, AI agents generate tests to build a “Golden master”
We then proceed to the code migration itself, with AI workflows respecting a depth-first traversal logic for each endpoint
Each agent iterates so that the migrated code exactly respects the Golden master tests
The different workflows are recombined to factorise the code and detect compilation errors
At the end, we obtain a report describing the bugs detected by the different agents and the tests we deliberately deviated from (because they were bugs we did not want to reproduce)
Guarantee identical behaviour
How to ensure the
accuracy of the migration ? This is the idea behind the method of the
Golden master. For the behaviour to be the same, each request must have the same response, even if it's a 500 error (internal server error). All responses, including bugs and errors, must persist in the same state, before and after the migration. These errors must be reported, but the correction of each must be controlled.
The change in behaviour must remain an exception. Therefore, it is necessary to write
a maximum number of tests : unit tests, integration tests, and end-to-end tests. Also test error handling and edge cases.
To strengthen the robustness and relevance of the tests, an “adversarial” structure : an attacking agent and a defending agent. The attacker only has access to the function contracts (what arguments does it take? what does it return? what is its purpose?). Its objective is to write tests: the happy paths, but also and especially the edge cases. In contrast, the defending agent is responsible for the migration itself. The defender will therefore need to iterate if necessary until the migration satisfies the tests generated by the attacker.
The idea behind this attacker/defender structure is to prevent the phenomenon of complacency (
sycophancy) of AI agents in test writing. By giving the attacker only the contracts and not the code of the functions it needs to test, we avoid tests being tautological and merely a retranscription of the code.
Writing tests is also an assurance for the future : in the event of future feature additions, robust and exhaustive tests already in place ensure the compatibility of modifications with the rest of the application.
Migrate its codebase
For the migration, the multi-agent system is architected as follows: an orchestrator generates sub-agents to migrate each of the endpoints. Its role is also to factorise the code and to pool the results at the end of the migration.
Each sub-agent is responsible for an endpoint. It must transcribe it into the new language. For example, at Galadrim, a model-view-controller (
MVC) in NET 4.8 was migrated to a .NET 9 backend and a React frontend.
Claude Code, from Anthropic, is an effective harness for this type of task: a main agent (the orchestrator) automatically generates the sub-agents responsible for test writing and migration. We proceed in waves: in each wave, 3 sub-agents are responsible for migrating endpoints. Each sub-agent is responsible for a group of endpoints (for example, four CRUD endpoints are part of the same group, which a sub-agent will have to migrate).
It may be necessary, to run an old codebase, to use a Windows VM (as is the case for .NET 4.8, for example) in order to generate end-to-end tests. The agents responsible for generating the tests then automatically connect to the VM to send the request and retrieve the legacy response to build the Golden Master.
Once all migration waves are completed, the orchestrator is responsible for piecing together the parts and factorising the code. This is particularly the case if functions were duplicated during the migration because they were necessary for several endpoints. It's also an opportunity to verify that there are no compilation issues (which simplification could cause) and that the migrated application launches correctly.
Detect bugs and security vulnerabilities
By scanning the test code, we can also detect major bugs previously unnoticed, and to security vulnerabilities. For example, we discovered in the legacy application that an endpoint was publicly exposed, without the need for authentication (unlike all others), and therefore the data it provided access to could be read without restriction. An exhaustive review of the code also allows for the detection of dead code, which would not be used and would pollute the rest. This can particularly manifest as tables present in the database but which are invisible to the application because they are unused.
In summary, migrating your codebase using a multi-agent system is, on the one hand, updating your code to ensure its maintainability and speed, but it's also about detecting, in the process,security risks and major bugs present in code that is no longer maintained.
Outcome: with AI, two weeks to migrate your codebase
How much would such a project cost? Manual migration would cost tens, even hundreds of thousands of euros and would take several months. This is why codebases do not evolve and remain in the same state long-term.
The multi-agent approach, on the contrary, drastically reduces the orders of magnitude. At Galadrim, we are converting a backend of 120,000 lines in two weeks. With Anthropic's Opus 4.8 model, the orders of magnitude are as follows:
2.5 billion tokens (mostly cache tokens, cheaper than input or output tokens)
1 800 $ of API cost
4,000 tests generated
40 hours of inference approximately, mostly parallelised (when 3 agents run for 1 hour in parallel, this represents 3 hours of inference for 1 hour of real time)