Back to Blog
In this article

API Architecture Mistakes That Quietly Break Back-End Scalability

Most back-end performance problems don’t appear suddenly. They are built in at the API design stage and stay unnoticed for months while the load is small. The system runs reliably for thousands of users and passes testing and investor demos. Then a marketing campaign, a seasonal peak, or a large enterprise client arrives, and response […]

API Architecture

Most back-end performance problems don’t appear suddenly. They are built in at the API design stage and stay unnoticed for months while the load is small. The system runs reliably for thousands of users and passes testing and investor demos. Then a marketing campaign, a seasonal peak, or a large enterprise client arrives, and response times grow several times over, the database hits its connection limit, and the cloud bill doubles. The cause is rarely “weak servers.” Most often it lies in architectural decisions that can’t be fixed by adding capacity and have to be reworked under load, when every hour of downtime costs money.

Why API Architecture Determines Back-End Scalability

An API is a contract between services, client applications, and external partners. The way this contract is designed determines how many requests a single user action generates, how much data travels over the network, whether you can launch ten more copies of a service, and how the system behaves when one of its parts slows down. Back-end scalability is the ability to handle growing load with a proportional, not exponential, increase in resources. If the API architecture doesn’t account for this, each new user costs more than the previous one.

For a business, API design mistakes have a very concrete price: product degradation at peak moments, rising infrastructure costs, slower development of new features, and the risk of losing customers exactly when the product starts to grow. The eight mistakes described below occur in projects of any scale, from startups at the MVP stage to enterprise systems with many years of history. They have one thing in common: at the start they seem like a convenient simplification, and under load they become a bottleneck.

Synchronous Call Chains: The Main Threat to Back-End Scaling

The move to microservices often comes down to replacing monolithic function calls with HTTP requests between services. The order service synchronously calls the inventory service, which calls the pricing service, which in turn calls the discount service. Each link adds network latency, and the response time of the entire chain equals the sum of the latencies of all participants. If each service responds in 100 ms, five sequential calls add up to half a second before the system has even executed its own business logic.

A more serious problem is cascading failures. When one service in the chain slows down, all upstream services hold open connections and threads while waiting for a response. Connection pools are exhausted, queues grow, and within a few minutes the problem of one link spreads to the entire system. A typical scenario: the recommendation service on an e-commerce platform starts responding slowly because of a heavy database query, and because of the synchronous dependency, checkout stops working, even though it is logically unrelated to recommendations.

The solution is to separate operations that truly require an immediate response from those that can be performed asynchronously through message queues or events. Comparing the two approaches helps determine where each one fits.

CriterionSynchronous interaction (REST/gRPC)Asynchronous interaction (queues, events)
Latency for the userSum of latencies of all services in the chainOnly the time to enqueue the task
Resilience to dependency failureFailure spreads to the calling serviceMessage waits in the queue until recovery
Scaling for peaksAll services must scale simultaneouslyThe queue smooths peaks, consumers scale separately
Data consistencyImmediateEventual consistency
Debugging complexityLower, simple request flowHigher, event tracing is required
When it fitsReading data, operations with an instant responseNotifications, reports, payment processing, integrations

The N+1 Problem and “Chatty” APIs as a Hidden Load Factor

The N+1 problem occurs when, to display a list of N items, the system executes one query for the list itself and N more queries for the details of each item. At the database level, this is a classic ORM mistake that is easy to miss during development: on test data with twenty records, the difference is invisible. In production with thousands of records, a single user request turns into thousands of database calls, and it is the first to appear in slow query reports.

API Architecture

The same pattern exists at the API architecture level. A “chatty” API forces the client to make many small requests to assemble one screen: the user profile separately, their orders separately, the status of each order separately. For a mobile app on a high-latency network, this means seconds of waiting, and for the back end, a multiple increase in the number of requests without any increase in the number of users. The opposite extreme is endpoints that return objects with dozens of nested fields, of which the client uses three, needlessly loading the network and serialization.

Practical solutions depend on the context. At the database level: batch loading of related entities and controlling the number of queries per HTTP call in tests. At the API level: aggregating endpoints for specific scenarios, the Backend for Frontend pattern for different clients, or GraphQL with a batch-loading mechanism and query complexity limits. It is important to note that choosing GraphQL does not by itself eliminate N+1, and without depth control it can even make the problem worse.

Missing Pagination: An API Design Mistake That Grows with Your Data

An endpoint that returns “all records” looks harmless at product launch: there are a hundred orders in the database, and the response takes a few kilobytes. A year later there are hundreds of thousands of records, and the same request loads megabytes of data into server memory, serializes it, and sends it to a client that displays only the first twenty rows. A few such requests at once can exhaust the service’s memory, and the endpoint itself becomes the easiest way to accidentally bring down the system with your own front end.

Even having pagination does not guarantee scalability. Classic offset pagination forces the database to read and discard all records before the requested page. Page 5 opens quickly, page 5000 slowly, and deep pages are precisely what integration and export scripts often request. Moreover, when new records are inserted between requests, items shift, and the client may receive duplicates or miss data.

For large and frequently updated datasets, cursor-based pagination is more efficient: the client passes the identifier of the last record received. The page size limit should be enforced on the server, not merely suggested by the client.

CriterionOffset paginationCursor-based pagination
Performance on deep pagesDegrades as the page number growsStable regardless of position
Stability when records are insertedDuplicates and gaps are possibleSequence is preserved
Jumping to an arbitrary pageSupportedNot supported
Implementation complexityMinimalRequires a stable sort key
Typical scenariosAdmin panels with small tablesFeeds, event logs, integration APIs

Server-Side State: A Barrier to Horizontal Back-End Scaling

Horizontal scaling, meaning launching additional copies of a service behind a load balancer, is the cheapest and fastest way to withstand growing load. But it works only when any copy can handle any request. If a service stores user sessions, shopping carts, intermediate results, or uploaded files in memory or on the local disk of a specific server, the next request from the same user must reach that exact server.

A popular “quick fix” is sticky sessions, where the load balancer pins a user to a single instance. The problem disappears temporarily, but new ones emerge: load is distributed unevenly, a restart or automatic scale-down drops sessions, and zero-downtime deployment becomes a risky operation. A typical example is a report generation service that stores the generated file locally and returns a link to it. While one instance is running, everything is fine; after a second one is added, some users get a 404 error.

The right approach is to make API services stateless. Sessions and temporary data are moved to a distributed cache, files to object storage, and authentication is built on tokens that any instance can verify. This API architecture makes it possible to use autoscaling in Kubernetes or cloud services without code changes and turns deploying new versions into a routine operation.

API Architecture

Ignoring Idempotency and Retries in API Architecture

In a distributed system, network failures are the norm, not the exception. A client sends a request to create a payment, the server processes it, but the response is lost due to a timeout. The client, having received no confirmation, repeats the request. If the API does not provide idempotency, the user is charged twice, the warehouse gets a double reservation, and support gets a complaint. Under load, the number of timeouts grows, and with it the number of duplicates.

The other side of the problem is so-called “retry storms.” When a service slows down, clients and neighboring services with aggressive retry logic start sending each request two or three times. The load on an already overloaded service multiplies, and it can no longer recover on its own. A system that would have withstood the peak goes down because of its own “reliability” mechanisms.

Rules that make retries safe for back-end scaling:

  • Idempotency keys for create and payment operations: the client passes a unique operation identifier, and the server returns the stored result instead of executing the operation again.
  • Exponential backoff with random jitter between retries, so that clients don’t send requests in synchronized waves.
  • A limited number of attempts and a retry budget at the service level, so that retries don’t exceed a defined share of traffic.
  • Correct response codes that clearly distinguish temporary errors worth retrying from permanent ones where retrying is pointless.

Missing Rate Limiting and Timeouts: When an API Doesn’t Protect Itself

An API without limits relies on all clients behaving correctly. In practice, one partner runs an integration script with no delays, one mobile app gets stuck in a request loop due to a bug, or a bot starts mass-querying the search endpoint. Without request rate limiting, such a client consumes resources meant for everyone, and service quality drops for every user.

Equally dangerous is the absence of timeouts on outgoing calls. HTTP client libraries often have very long or infinite default timeouts. If an external payment provider or geolocation service stops responding, server threads hang while waiting, and within a few minutes the API stops accepting new requests, even though its own code is completely sound.

Basic protective mechanisms that should be built into the API architecture from the first version:

  • Rate limiting at the API gateway level, with limits per client or key and a 429 response with a header indicating when to retry.
  • Explicit timeouts on every outgoing call, aligned with the overall response time budget.
  • Circuit breaker, which temporarily stops calls to a dependency that consistently fails to respond and returns a fallback response.
  • Resource isolation (bulkhead): separate connection pools for different dependencies, so that the failure of one doesn’t exhaust resources for all.
  • Request size control: limits on request body size, the number of items in batch operations, and query complexity.

An Ineffective Caching Strategy as a Scalability Constraint

Caching is the most powerful tool for scaling a back end under read-heavy load, and at the same time a source of some of the hardest problems to diagnose. A common mistake is caching only at the application level and ignoring HTTP caching. When an API doesn’t return Cache-Control and ETag headers, CDNs and clients can’t reuse responses, and every request reaches the back end, even if the data hasn’t changed for hours.

The second typical problem is the simultaneous expiration of a popular key. When the cache for the catalog home page becomes invalid at a moment of high traffic, hundreds of requests hit the database at once to rebuild the same value. During a sale on a marketplace, this is enough to make the database stop responding for several minutes. It is countered by locking while the value is recomputed, background cache refresh before expiration, and random TTL jitter.

The third mistake is caching without well-thought-out invalidation. A team adds a cache to speed up the API and then finds that clients see outdated prices or stock levels. In response, the TTL is cut to a few seconds, and the cache effectively stops working. The caching strategy should be defined together with the API contract: which data can be stale and by how much, which events should reset the cache, and where personalized responses are stored.

An API Without Observability: Scaling the Back End Blind

A team that can’t see how the API behaves under load learns about problems from users. Average response time is a common but misleading metric: it can stay within normal limits while every twentieth user waits five seconds. Without distributed tracing, it’s impossible to understand which service in the call chain is responsible for the latency, and optimization turns into guesswork.

Key metrics worth collecting for every critical endpoint:

  1. p95 and p99 latency percentiles, which show the experience of the slowest requests rather than the average picture.
  2. Error rate, broken down by response codes and client types.
  3. Throughput, the number of requests per second per endpoint.
  4. Resource saturation: utilization of connection pools, queues, memory, and CPU.
  5. Number of database queries per HTTP call, an early indicator of N+1 and inefficient queries.

Observability also has a financial effect. When it’s clear which endpoints generate the most load, the team optimizes exactly those rather than scaling the entire infrastructure. Load testing with realistic scenarios before a release or marketing campaign reveals the system’s limit in advance, when there is still time to push it back, rather than in the middle of a traffic peak.

Scalability Is Built into the Contract, Not the Servers

Most back-end scaling problems can’t be solved by buying more powerful servers. Synchronous chains, N+1, unbounded responses, server-side state, unsafe retries, missing protective mechanisms, chaotic caching, and an opaque system are all architectural decisions built into the API contract. Each of them on its own seems minor, but together they determine whether the product will grow along with its audience or spend more and more resources maintaining what already exists.

API design mistakes are cheapest to fix before clients, partners, and mobile apps depend on them. For products that are already live, an effective first step is an independent API architecture audit and load testing with a team experienced in scaling high-load systems. Such an audit reveals bottlenecks before users discover them and makes it possible to plan fixes in stages, without halting development and without risk to the business at peak moments.

About the author

Stanislav N.

Stanislav N.

Senior WordPress / PHP Developer

Stanislav builds custom WordPress and PHP systems at Meduzzen, from bespoke themes and ACF-driven blocks to multilingual WooCommerce stores. His work includes an SEO-optimized microsite template system for an AdTech lead-generation client and the multilingual e-commerce platform behind a German confectionery brand, with real-time cart updates, a loyalty program, and Stripe, PayPal, and Google Pay checkout. He builds sites that stay fast and maintainable long after launch, not just on day one.

Have questions for Stanislav?
Let’s Talk

Read next

You may also like

Quick Chat
AI Assistant