Talk to Us +91 98847 45599

Technology & Engineering

Why Good Software Architecture Is Mostly About Decisions

The boxes on an architecture diagram rarely cause trouble. The decisions behind the lines do, and most teams never write them down.

Vianmax Editorial9 min read

Technical illustration of a path that branches at several decision points. One branch continues in blue while the alternatives end in marked terminations.

Anyone who has inherited a system knows the feeling. There is a diagram, often a good one: services in tidy rectangles, arrows labelled with protocols, a database drawn as a cylinder because databases are always drawn as cylinders. It tells you what exists. It almost never tells you why.

Why does the reporting service read from a replica instead of the primary? Why is there a queue between these two components and not the other two? Why did someone build a custom scheduler when the platform already had one? The answers existed once, in someone's head or in a meeting nobody minuted. Now the system works, mostly, and nobody is confident enough to change the parts they don't understand.

This is the quiet failure mode of software architecture. It isn't bad design. It is design whose reasons have evaporated. And it suggests that the useful definition of architecture is not "the structure of the system" but something closer to this: the set of decisions that are expensive to change, and the reasons they were made.

The diagram is the least important artefact

Diagrams are useful. They are how teams build a shared picture quickly, and a system nobody can sketch is usually a system nobody fully understands. But a diagram is a snapshot of outcomes. Two teams can draw identical boxes for very different reasons, and those reasons determine what happens when requirements move.

Consider a general example. Two systems both place a message queue between order intake and order processing. In the first, the queue exists because processing is slow and intake must stay responsive during bursts. In the second, it exists because the processing team deploys on a different schedule and wanted to be able to take their service down without losing requests. Same box, same arrow. Now imagine processing gets fast enough that the burst problem disappears. The first team can reasonably consider removing the queue. The second cannot, because their reason still holds. The diagram cannot tell them apart.

So when we review an architecture, the diagram is the place to start the conversation, not the thing being reviewed. The questions that matter sit behind it. What was this optimised for? What was traded away? What would have to change for this to be the wrong choice?

Reversibility is the real axis

Not all decisions deserve the same care, and treating them as if they do is its own mistake. Teams that deliberate over every library choice move slowly and still get the big things wrong, because their attention is spread evenly across problems of very uneven weight.

The most useful single question about a decision is how expensive it would be to reverse. Choosing a date-formatting library is cheap to undo. Choosing how customer identity works across your products is not. Choosing a web framework sits somewhere in between, cheap in week one and very expensive in year three.

A few categories of decision tend to be expensive to reverse almost everywhere:

  • how data is shaped and who owns it
  • the contracts between components, especially ones other teams or external parties depend on
  • identity, permissions and tenancy models
  • anything that determines where state lives and how it is kept consistent

Everything else deserves a lighter touch. Make the call, note it, move on. Spend the deliberation where reversal is costly, and spend it early, because the cost of reversal almost always rises with time.

Architecture is not the structure of a system. It is the set of decisions that are expensive to change, and the reasons they were made.

Simplicity is a decision, and so is flexibility

There is a common instinct, particularly among capable engineers, to build for flexibility. Make the storage layer pluggable in case we change databases. Add a plugin system in case customers want extensions. Abstract the payment provider in case we switch.

Some of these turn out to be wise. Many don't. The problem is not that flexibility is bad. It's that flexibility has a price, paid immediately and continuously, while its benefit is speculative and deferred. Every abstraction is another layer someone must understand before making a change. Every extension point is a contract that must be kept stable. An interface designed for three hypothetical implementations tends to fit the one real implementation badly.

The honest framing is that simplicity and flexibility are both decisions, and both have costs. Choosing the simple path means accepting that some future changes will be more expensive. Choosing the flexible path means accepting that every change today is a little more expensive, and that the flexibility might never be used. Neither is a default. What matters is making the trade consciously, with some evidence about which futures are likely.

One heuristic that holds up well: build the abstraction when you have the second concrete case, not before. With one case you are guessing at the shape of the variation. With two, you can see it.

The lines matter more than the boxes

When systems fail in ways that surprise their builders, the cause is rarely inside a single component. It is usually at a boundary. Two services disagree about what a field means. A downstream system retries a request that was not safe to retry. A third-party API changes its error format. A timeout on one side is longer than the timeout on the other, so requests are abandoned and then completed anyway.

Two kinds of boundary deserve particular attention.

Data boundaries decide which component owns which facts. When two parts of a system can both write the same data, you have a consistency problem, whether or not you have noticed it yet. When a component reads data it doesn't own, you have a coupling that will surface the next time the owner changes its schema. Data also tends to outlive the code around it. Applications get rewritten; the tables usually survive, carrying every early modelling decision with them. This is why data modelling deserves more care than almost anything else in the early weeks of a project.

Integration boundaries are the points where your system meets something you don't control: a payment gateway, a broker, a government service, a partner's API. These deserve a different posture. Assume the other side will be slow, will occasionally be wrong, will change without notice, and will fail at the worst moment. The design question is not whether it will fail but what your system does when it does. A common and effective pattern is a dedicated adapter per external system, so that the rest of the codebase speaks one internal language and the peculiarities of each provider are contained in one place. It is the approach we use for broker connectivity in AlphaSync, where each broker's order types and error responses differ and the strategy code should not have to care.

Deciding what the system will be able to tell you

Observability is usually discussed as an operational concern, something to add once a system is running. That framing undersells it. What a system can tell you about itself is decided by its architecture, and it is hard to retrofit.

If requests do not carry an identifier from the edge through every internal call, you cannot trace one user's problem through the system later. If state changes are not recorded as events with timestamps and causes, you can see that a value is wrong but not how it became wrong. If components do not distinguish between "I failed" and "someone I depend on failed", every alert looks the same at two in the morning.

A useful exercise during design is to walk through the likely failures and ask, for each one: how would we know, how quickly, and what would we look at to understand it? The answers often change the design. They might add a correlation ID, an audit table, a health signal that reflects real work rather than a process being alive. These are cheap to build in at the start and surprisingly expensive to add once the system is carrying load.

Technical debt is a decision you forgot you made

"Technical debt" has become a catch-all for any code people dislike, which makes it less useful than it should be. The original metaphor was more precise: you take on debt deliberately, to gain speed now, knowing you will pay interest until you repay it.

Understood that way, debt is not a failure. Shipping a narrower solution to meet a real deadline is often correct. The trouble starts when the decision is not recorded. Six months later, nobody remembers that the single-region deployment was a conscious shortcut with a plan to revisit it at a certain scale. It simply looks like how things are. The debt is still there, still accruing interest, but it has stopped being a decision and become an assumption.

Debt that is written down (what we chose, why, what it costs, when we should revisit it) can be managed. Debt that isn't tends to be discovered during an incident.

Writing decisions down

If architecture is mostly decisions, the most valuable architecture document is a record of them. Many teams have adopted some form of lightweight decision record: a short document per significant decision, kept with the code, written at the time the decision is made. The format matters less than the habit, but the useful ones tend to capture the same few things.

Contextforces, constraints01Optionsand why they lost02Decisionstated plainly03Consequenceseasier / harder04Review triggerwhen to reopen05When the trigger condition is met, the decision is reopened with fresh context
Anatomy of a lightweight decision record. The review trigger turns a permanent fact into a considered position with an expiry condition.

The context explains the forces at play: requirements, constraints, what was known at the time. The options show that alternatives were considered and why they lost, which saves future teams from re-litigating settled questions. The decision is stated plainly. The consequences are the part most often skipped and most valuable later: what this makes easier, what it makes harder, what we are now committed to. And a review trigger names the condition under which the decision should be reopened, such as a traffic level, a new regulatory requirement or a second customer type.

That last element is the one we would add if a team has only one habit to adopt. It turns a decision from a permanent fact into a considered position with an expiry condition. It also gives the next engineer permission to change it, which is often what they need most.

These records are short, often a page. They do not replace diagrams or specifications. They explain them. In our own work, the architecture phase of every engagement produces documented and agreed system architecture, data models and API contracts before the full build starts (our approach describes the phases), precisely so that the reasons survive the people who were in the room.

What good looks like

A well-architected system is not one where every decision was right. That standard is impossible, because decisions are made with incomplete information and circumstances change. A well-architected system is one where the important decisions were made deliberately, their trade-offs were understood, the expensive-to-reverse ones received proportionate care, and the reasoning was preserved well enough that someone new can judge whether it still applies.

That is a less glamorous definition than the one implied by clever diagrams. It is also more useful. It suggests that the best thing an architect can leave behind is not a picture of the system but an honest account of the choices that shaped it, including the ones they were unsure about.

The technology itself matters too, of course. There is a reason we keep our core stack deliberately small: fewer tools, understood deeply, mean fewer decisions made by default and more made on purpose. But the stack is downstream of the habit. Teams that decide well, and write down why, tend to end up with systems that can be changed. That, more than any particular pattern, is what good architecture buys.

Filed under Technology & Engineering · Vianmax Editorial ·

Examples in this article are general and conceptual unless stated otherwise. Where Vianmax products or work are mentioned, the description matches what is published elsewhere on this site.

Keep reading

Product & Digital Systems

Software Products Are Systems of Experience

A user's opinion of a product is formed as much by a slow search, a confusing permission error or a late email as by anything drawn in a design file.

8 min read

How we architect before we build

Every Vianmax engagement documents architecture, data models and API contracts, and agrees them, before the full build starts.