The AI plateau in architecture
What makes architecture worth defending?
Currently architecture is rapidly changing with a snow-ball-like effect taking place, it is a change which we cannot stop nor can we predict. And in some cases it feels like it is an attack on the profession, and our need to remain human-centric. And so, the question is, what makes architecture worth defending?
The AI plateau in architectural practice
Every architecture firm right now is running the same experiment, whether they’ve named it or not. Someone on the team is quietly testing an AI tool on a real project, and the results are a combination of genuine wins and quiet disasters. A rendering that would’ve taken a day now takes ten minutes. A zoning check that looked airtight turns out to have missed a setback nobody caught until it was expensive. Openly, very few are tracking any of this systematically, so the story studio’s tell themselves is whatever anecdote happened last week.
That’s the plateau. Not a slowdown in the technology, but a stall in our ability to say, with any precision, where it helps and where it quietly costs us more than it saves.
The instinct is to skip straight to the interesting part: what should we do about AI right now? We think that question is premature. Before you can answer it, you need to know what kind of tool you’re holding, and that means understanding what AI has and hasn’t been good at since the day someone first gave it a name.
Before we can talk about the plateau, we need to take a step back and ask how we got here.
Defining AI
“Artificial intelligence” first showed up in 1955, when John McCarthy coined it ahead of the Dartmouth Workshop that ran through 1956 (Mucci, Finio, & Downie, 2026). That workshop is generally treated as the birthplace of AI as a field.
The dawn of machine logic (1950s–1960s)
After Dartmouth, the field rode a wave of optimism. Early researchers tried to teach machines to mimic human problem-solving and basic learning.

The shift to specialized expertise (1970s–1980s)
Once it became clear that replicating general human intuition was much harder than expected, the field narrowed its ambitions. Instead of building broad “thinking machines,” researchers built expert systems.
Stanford’s DENDRAL is the classic example: a program that automated complex decision-making in organic chemistry, matching human experts within a narrow domain. This era proved something useful, general intelligence was a hard problem, but targeted computational help worked well. It’s a similar arc to how architectural tools themselves evolved, from basic drafting up to full BIM software.
The modern resurgence and where we stand now
Over the past couple of decades, more computing power and more data let the old neural network theories finally work. The field moved away from rigid, rule-based programming toward deep learning, where systems find patterns and generate outputs on their own.
That’s what brings us to the plateau we want to address. We now have generative and analytical tools that are powerful, and we’re also starting to see their limits clearly: spatial intuition, physical context, human-centered design logic. The pattern across all three eras – logic theorists, expert systems, modern pattern recognition – is the same. AI has always been better at specialized, bounded tasks than at holistic human judgment.
What is the new balance between artificial and human intelligence in AEC?
Most of the public debate is about whether AI will replace practitioners. We think that’s the wrong question. The harder, more useful one is narrower: which parts of your core work get better when a machine does them, which parts fall apart, and where’s the line between the two once the novelty wears off?
That’s what we are trying to figure out at D/DOCK. We think this grey zone only becomes visible once the novelty fades, so mapping it is the point of this pilot.
Part of this six month journey we will explore AI in architecture through various lens through a three phased approach.
AI is currently strong at bounded, specialized data problems. Whether it can handle holistic, long-horizon architectural work is still an open question. To test that, we’ll deploy, benchmark, and evaluate both agentic and non-agentic workflows across the three phases below.
Phase one: baseline (months 1–2)
We’ll start by mapping current workflows and setting baseline metrics. (Kokotajlo, Alexander, Larsen, Lifland, & Dean, 2025), who argue AI performance needs to be measured against clear human benchmarks. That’s the approach of this phase. First, we isolate bounded, specialised tasks where an AI agent could plausibly help without needing full spatial intuition: building code compliance checks, early massing iterations, compiling project specs and generative media. We log how many hours our team currently spends on each. At the same time, we identify and procure the tools we need for the pilot. The most capable autonomous tools will require real budget, and this data will form part of the spine for this research.
Phase two: supervised agent testing (months 3–4)
Here we test whether the AI can act as an autonomous researcher and technical assistant, not just a drafting tool. We give it complex, multi-step instructions. Reviewing a local zoning code and applying setback requirements directly to a site boundary, for example, to see whether it can handle compound logic. This phase is about documenting the “stumbling” period of adoption. D/DOCK’ers review every output and log spatial errors, physical impossibilities, and misreadings of the brief. That’s the clearest way to see where machine intuition still breaks down.
Phase three: long-horizon benchmarking and ROI (months 5–6)
The last phase asks whether AI can sustain context over time, not just execute one task well. We stress-test that by having it maintain something like a schedule of materials autonomously while a design goes through iterative changes over several weeks. Then we compare completion times against our phase-one human baselines. The number that matters most here is what we’re calling the “editing tax”: whether the time AI saves on the first draft gets eaten up by the time senior employee’s spend fixing its mistakes.
While we know that there are more complex ways to investigate AI and it’s role, we are also aware of our own limitations. As our knowledge base matures, so does the complexity of our questions. I do not believe that we will ever be able to catch up with AI, and so at this point in time we are making guesses about the strategy of AI.
Bibliography
Mucci, T., Finio, M., & Downie, A. (2026, August 19). IBM. Retrieved from IBM.com: https://www.ibm.com/think/topics/history-of-artificial-intelligence
Kokotajlo, D., Alexander, S., Larsen, T., Lifland, E., & Dean, R. (2025, April 3). Mid 2025: Stumbling Agents. Retrieved from AI 2027: https://ai-2027.com/
Contact
Every design starts with a question. We’re ready for yours. Let’s explore, push boundaries and bring people together. We’d love to hear your thoughts, so send us a message.