OTWopentechwire
Tech Intelligence, Openly Wired
Dev

315 Pull Requests, Zero Lines Typed: How a Mews Product Builder Runs a Team of AI Agents in Production

Ettore Zotarelli trained in hotel management and cannot write code by hand. He ships to a live platform used by 15,000 hospitality businesses. In a written interview he lays out the workflow, the review rules and the failure modes, and answers the strongest objections to the model.

RN
Ryan Ng
·
Sep 30, 2026
15 min read
A quote card reading "Technology is most powerful when it makes people's lives easier, so they can focus on what truly matters: hospitality," attributed to Ettore Zotarelli, Lead Product Builder at Mews, beside a photo of a smiling bearded man in a navy suit with his arms crossed.
Ettore Zotarelli, Lead Product Builder at Mews.Credit: Courtesy of Ettore Zotarelli. Graphic: OpenTechWire.

Most writing about AI-native software development comes from people without a production system, paying customers or an on-call rotation. Ettore Zotarelli has all three, which makes his account unusually useful.

Zotarelli is Lead Product Builder at Mews, the Amsterdam-based hospitality platform that raised $300 million in January 2026 at a $2.5 billion valuation and reports 15,000 customers in 85 countries. Its software handles reservations, check-in, payments and guest profiles for working hotels. The Product Builder role collapses product management, design and engineering into one person who directs AI coding agents and owns the result. Mews piloted it in one team in mid-2025. According to Zotarelli, a July 2026 reorganization of research and development replaced the classic squad with Product Builders in several parts of the company. That same month Mews cut about 15% of its roughly 1,350 staff, according to Skift and Hotel Dive, as part of what founder Richard Valtr described as a shift to an AI-native company.

The model has critics inside the industry and questions inside Mews itself. In a May 2026 post on the company's developer blog, fellow Lead Product Builder Quintin Botes listed slow pull request turnaround, ownership ambiguity and the risk that builders become permanent maintenance teams for everything they ship. Google's 2025 DORA report, drawing on nearly 5,000 technology professionals, found that 90% now use AI at work, that about 30% have little or no trust in AI-generated code, and that AI adoption still correlates with lower delivery stability.

Zotarelli's numbers, drawn from Mews' GitHub history and reported by the company, show what one builder's output looks like.

Table with bars showing merged pull requests by month at Mews, March to 20 August 2026: March 20, April 30, May 47, June 94 (the peak), July 82 and August 42 up to the 20th, 315 in total. Lines added rose from 311 in March to 17,201 in August, 50,371 in total, and lines removed totalled 8,050.

That is 315 merged pull requests and roughly 50,000 lines added. The count peaked in June while the size of each change kept growing. "I got better at scoping work for agents, not just at running more of them," he said.

What follows is an edited selection of his written answers. The interview has been condensed, and spelling has been adapted to American style.

The role

In one sentence, what are you accountable for that a product manager is not?

Outcomes and the code that produces them. A product manager is accountable for deciding what gets built and why, then handing that decision to other people to design and implement. I am accountable for the decision, the design, the shipped implementation, and what happens when it breaks.

You have said this job could not have existed before. What changed?

Several things changed and the role exists because they arrived together. Agents got good enough to hold architectural context across a whole repository rather than a single file, so the unit of work became a change rather than a function. They moved into CI, so they run asynchronously against a branch instead of waiting on someone in a terminal. Reusable instructions became a thing you can write once and have every agent read, which is how local convention gets enforced without a person enforcing it. And the tooling became good enough to work across several repositories together, which is where most real product work actually lives rather than a single codebase.

That last one mattered most for me. A single-repository assistant is a productivity tool. Something that can carry a change coherently across the frontend, the backend and the services in between is what makes owning a whole product feasible for one person.

The workflow

Walk us through one real task from the last two weeks.

Migrating ten backend-rendered settings pages to the frontend. That meant changing programming language, moving onto our design system, and working across the frontend, the backend and the services in between at the same time.

The agents did the mechanical work: translating logic, mapping old markup onto design system components, wiring localization, generating tests. That is high volume, well defined, and exactly what a migration skill is written for.

The rest was mine, and it splits into three kinds of decision. Product: I did not port those pages faithfully. Two of them became one, because the split existed for a technical reason that no longer applied. I added bulk editing and bulk deleting where the data made it obviously useful, and left it out where it did not. Design: which components to use, and what the experience should be for someone configuring their property for the first time. Engineering: reviewing every diff, and stepping in on state handling and the edge cases the old pages handled implicitly, where the agent produced code that looked correct and quietly lost a behavior nobody had documented.

How many agents are actually running?

I usually have around ten sessions running in parallel, each scoped to a project. Each dispatches its own subagents as needed, sometimes twenty or more inside a single job, for things like reading prior code and researching documentation, drafting the changes, writing tests, checking localization.

Separately, we have a CI-based agent that reviews every pull request across all our repositories. It runs automatically. I do not invoke it and I cannot skip it. It comments, and it fixes small things itself.

The practical consequence is that my loop is asynchronous. I prompt across several streams, then come back thirty to forty-five minutes later to review. The constraint on my day is not how fast the agents produce. It is how fast I can read and judge what comes back.

What does correcting an agent look like?

If I only fix the diff, the same mistake returns next week in a different file. So correction means editing the instructions. Our repositories carry agent rules and skills that encode local convention: how design system components are used, how tests are structured, how localization works. When a review round surfaces a repeated mistake, I write that learning into those rules.

The bigger change is how code gets written in the first place. We now write for agents as much as for people, which means more context and more comments explaining why. If nothing in the code explains why this is a button rather than a toggle, or why a pattern was deliberately broken for one unusual case, the agent has no way to know, and it will confidently normalize the exception away. Intent used to live in the heads of the people who made the decision. It now has to live in the file.

Review and guardrails

What is your review discipline before anything reaches production?

Nothing reaches production today without a human.

The rules: any pull request carries a named human author who owns the change as if they had typed it. A pull request authored by an agent needs two named humans, one who authorizes it and a different one who approves it, and neither can be an agent. Every pull request gets an AI review before a human is asked to look, so human review time goes on whether this is the right change rather than on missing tests and convention breaches.

Around that sit the other safeguards: comprehensive automated testing in CI, feature flags on effectively everything so a bad change is switched off rather than rolled back under pressure, and staged rollout.

We track how much human correction a pull request actually needs, and plenty already pass human review with no comments at all. There will always be sensitive areas, architecture in particular, that go through a human regardless of how good the AI review is. If eighty percent of changes could be safely cleared by AI review, that is already an enormous productivity gain, and the remaining twenty percent is exactly where you want your senior people spending their attention anyway.

What do agents get confidently wrong?

The pattern I watch for is confident reinvention. An agent writes something locally correct, that passes tests and reads well, while ignoring that the codebase already solves that problem somewhere else. You end up with two ways of doing one thing, which is how a codebase gets expensive.

The second failure mode is imitating the shape of a pattern without its constraints. The code looks like the code around it, but the assumption that made the original safe is missing.

Both are hard to catch from the diff, because the diff looks good. Part of the answer is prompting for it explicitly: "check how this other, similar page does it first", "make sure this pattern does not already exist somewhere else before you create a new one", "find the prior art and tell me why you are not using it". Asked that way, the agent usually finds the existing solution itself. Not asked, it will happily build you a second one.

The hard questions

Critics argue AI produces code faster than organizations can understand it, and that the real risk is architectural entropy. Your answer?

It is a real risk and I would not argue against it in general terms. I have seen the shape of it: a change of a few hundred lines, rated low risk, that broke a build pipeline because nobody fully understood the surface it touched. It was caught in CI and never reached a customer.

What I would push back on is the idea that this is new. Undocumented architectural knowledge living in a few senior engineers' heads has always been the failure mode. An agent cannot read the intent in somebody's head. It only reads what is in the repository. So it forces the discipline we always claimed to want: write down why, document the decision, explain the exception.

The test I hold myself to is whether a human can still explain the system. That is not automatic, and I will not claim it is settled. 2026 is the first year anyone has a codebase built this way, so the two-year question is genuinely open.

You do not come from engineering. Where is the honest gap?

I can read code. I can follow a pull request, understand what a change does and judge whether it belongs. I cannot explain it to you line by line.

The bigger gap is domain rather than syntax. I am very good at extending and improving things that already exist. When something is genuinely new, either new inside the system I know, or new because it is a system I am not familiar with, I need an expert before I start. Not at review, at planning. I book half an hour with someone who owns that code, walk them through the plan, and let them tell me how that code is actually written, what the underlying constraints are, and what to watch out for.

The failure mode I am guarding against is not writing bad code, it is confidently building the wrong thing in a domain where I could not tell.

Mews runs live hotel operations. When agent-assisted code causes an incident at two in the morning, who is accountable, and has it happened?

It has not happened to me, and that is a statement about my own work rather than a claim about the industry.

We have twenty-four hour on-call coverage, with named people responsible for specific parts of the product. I might not be the one awake at two in the morning, and I would not recommend a model where I was. Someone is, and that person can turn off a feature flag and make the changes needed to stabilize things.

Accountability itself is unambiguous. The named author of a pull request owns the change as though they had written every line by hand. When production breaks, an incident lead owns coordination and resolution. An agent is never the accountable party, deliberately, because accountability that cannot be assigned to a person is not accountability.

Guest data is regulated. How do you keep personal information out of an agent's context window?

On the building side, the control is that coding agents have no path to production data. They work against source code and synthetic or anonymized data. That is a written rule rather than a convention: never connect AI tools to production data. Anything touching live guest state, a migration, a backfill or a deletion has to be approved and executed by a human. An agent can draft the plan and cannot run it. Agent configuration is managed centrally rather than left to each engineer's laptop, and it blocks credential and secret access.

Beyond that I would rather not characterize our formal commitments from memory, because this is exactly the area where an approximate answer is worse than none. Mews publishes its security and compliance documentation at trust.mews.com, and our security team is the right source for specifics.

If a builder owns everything he ships indefinitely, he becomes a maintenance team of one. Have you hit that wall?

It is the right critique and the one I think about most. I have not hit the wall, and I am early enough that you should treat that as early data rather than a refutation.

We build independently, we do not maintain in isolation. Several people know each product well enough to pick it up, on-call responsibility is shared, and I have a backup builder covering my products when I am away.

Maintenance itself got dramatically cheaper. Refactors that used to take weeks happen in hours, and instead of one page at a time I refactor ten in parallel. The structural answer is still that ownership is scoped to a domain rather than accumulating project by project.

The same critique says the surrounding organization is still built for squads. Is that true at Mews?

Partly true, and it was more true a year ago. What is still true: the moment a change crosses into someone else's domain, I move at the speed of their queue. That is not an AI problem and no amount of agent throughput fixes it. Fast individuals inside a slow dependency graph still wait.

Anyone claiming they have removed cross-team latency has probably just stopped measuring it.

What the shift is really about

What do people get most wrong?

That it is about speed. Producing code was never the most expensive part of software, deciding correctly what to build was. But code used to be expensive enough to distort every decision around it, and that is what actually changed.

Now I can build a working version, show it to stakeholders, and have them tell me it is wrong or complicated enough that we should drop it. What I wasted was a few prompts. Being wrong got cheap, which means we can afford to find out rather than argue about it.

So the volume goes up and the cost of a bad foundational decision goes up with it. A wrong architectural call used to be limited by how fast people could implement it, and now it propagates through ten pull requests before anyone notices. The scarce skills are judgment, review and taste, and they are all harder to hire for and slower to build than the ones this shift made abundant.

Does "Product Builder" survive as a title?

I think it stays, though the label may change into some other fashionable phrase. A Product Builder does not replace specialists. It puts more pressure on them, and raises the value of the very best ones. When I ask an agent to build a page, it is accurate because there is an exceptionally well-written design system behind it that spells out the do's and don'ts of every component.

This role only works on top of excellent architecture, an excellent design system, and the specialists who build and maintain them. The people writing the rules matter more now, not less, because everything I produce inherits their judgment at scale.

What should a product manager do in the next six months?

Ship something end-to-end yourself. One real thing, in production, that you own including the parts you find uncomfortable.

Learn to read differences rather than code. I do not read a pull request line by line and I do not need to. I read what changed: what was there before, what is there now, what was removed. That skill is learnable without learning to program, and it is the skill that actually gates everything else.

Build reusable instructions rather than one-off prompts, and read them. At the end of a long session with a lot of back and forth, ask the agent what it learned that would be worth repeating, and have it write that up as a reusable skill.

Get as close to on-call as your organization allows. Owning the consequences changes your judgment faster than any course.

And pick a real system to learn on. Anyone can generate an app nobody uses and that nobody would ever pay for. The learning is in shipping to people who will notice when you are wrong.

What to take from it

Read as an operating manual, Zotarelli's account reduces to five practices that any engineering organization can test. Treat instructions, not diffs, as the thing you fix. Put an automated reviewer in CI that nobody can skip, and reserve human review for the question of whether a change is right. Require a named human owner on every change and two on anything an agent authored. Ship behind feature flags so recovery is a switch rather than a rollback. Bring in the domain owner at planning, not at review.

Read as evidence, it has limits, and he states most of them himself. The figures are self-reported by one builder at one company and have not been audited. The role sits on top of a mature design system, extensive automated testing and senior specialists, which many organizations lack. The long-term maintainability of a codebase built this way is unknown. And the context is a company that reduced headcount while adopting the model, which means the productivity claim will ultimately be judged on incident rates and customer outcomes rather than on pull request counts.

For engineering leaders, the more durable point may be the one about where cost has moved. If agents can produce a correct-looking change in minutes, the scarce resources are the reading time of people who can tell correct from correct-looking, and the written intent that lets an agent tell a rule from an exception. Teams that invest there will get the benefit. Teams that only add agents will get more code.

Read next
Dev

Trust Fractures as Beijing AI Startup Silently Harvests Developer Code

Marcus Halloran · 5 min
Dev

Vibe Coding Splits Developers as LLMs Write Nearly Half of Production Code

Linh T. Pham · 5 min
Dev

Android Developers Face New Memory Caps as AI Data Centers Drain Chip Supply

Arjun S. Mehta · 5 min
Spot something wrong? Email corrections@opentechwire.com. We log every correction publicly.