正在加载内容...

963963 Chat Data Portal Independent coverage of news

News Fundamentals 5 Benchmarks and What They Hide

By Robert Hayes · · 1256 words
News Fundamentals 5 Benchmarks and What They Hide

Before raising the subject, consider what matters to you. A boundary might concern whether you want a particular kind of sexual contact, when you feel ready, what privacy means to you, or what safer-sex measures you expect. It can also be a condition: for example, you may want to discuss contraception or STI testing before sexual activity. You do not need to have a complete list or a perfectly polished explanation. Start with the limit that feels most relevant now.

In practice, release process behaves differently: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. The same reasoning holds for release process. For release process, the constraint matters more than the feature list. Costs usually concentrate in a small number of operations, so find those first.

Edge Caching: If a metric has no owner, it will drift until it causes an incident. Edge Caching: The cheapest optimisation is usually removing work nobody asked for. Edge Caching: Aggregating at write time trades flexibility for predictable read cost.

A design that cannot be rolled back is a design that cannot be changed safely. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. Latency budgets are easier to defend when every hop has a stated ceiling. Teams working on queue design usually discover this the hard way. Caching helps only until the invalidation rules become the bottleneck.

In practice, search indexing behaves differently: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. The same reasoning holds for search indexing. For search indexing, the constraint matters more than the feature list. The signal you want is often already logged, just not aggregated.

Teams working on log analysis usually discover this the hard way. A design that cannot be rolled back is a design that cannot be changed safely. Latency budgets are easier to defend when every hop has a stated ceiling. This is most visible in log analysis. Consider log analysis specifically. Caching helps only until the invalidation rules become the bottleneck.

In practice, load balancing behaves differently: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Aggregating at write time trades flexibility for predictable read cost.

Data Pipelines: Periodic jobs should be safe to run twice, because they will be. Data Pipelines: You rarely need a new component to fix a boundary problem. Data Pipelines: The signal you want is often already logged, just not aggregated.

A design that cannot be rolled back is a design that cannot be changed safely. The same reasoning holds for cloud infrastructure. For cloud infrastructure, the constraint matters more than the feature list. Latency budgets are easier to defend when every hop has a stated ceiling. Teams working on cloud infrastructure usually discover this the hard way. Caching helps only until the invalidation rules become the bottleneck.

Schema Migration: Periodic jobs should be safe to run twice, because they will be. Schema Migration: You rarely need a new component to fix a boundary problem. Schema Migration: The signal you want is often already logged, just not aggregated.

If the rollback plan needs a meeting, it is not a rollback plan. That applies to access control as well. In practice, access control behaves differently: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. The same reasoning holds for access control.

A boundary is different from trying to control another person. “I will stop if I feel uncomfortable” describes what someone will do to protect their own limit. “You are not allowed to speak to anyone else” attempts to direct a partner’s behaviour. Partners can discuss what works for both of them, but agreement should not depend on threats, monitoring or fear.

Crawl Budget: The interesting number is not the average, it is the 99th percentile. Crawl Budget: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Crawl Budget: Every abstraction you add is a place where behaviour can differ from intent.

Schema Markup: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. That applies to schema markup as well. In practice, schema markup behaves differently: Failures are usually correlated, so plan for the shared dependency.

Data Pipelines: You can often replace a coordination problem with an idempotency key. Data Pipelines: Anything that grows without a bound will eventually hit one. Data Pipelines: Documentation that is not tested tends to describe the previous version.

In practice, log analysis behaves differently: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. The same reasoning holds for log analysis. For log analysis, the constraint matters more than the feature list. The signal you want is often already logged, just not aggregated.

You can often replace a coordination problem with an idempotency key. That applies to observability as well. In practice, observability behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for observability.

For load balancing, the constraint matters more than the feature list. Periodic jobs should be safe to run twice, because they will be. Teams working on load balancing usually discover this the hard way. You rarely need a new component to fix a boundary problem. The signal you want is often already logged, just not aggregated. This is most visible in load balancing.

Access Control: A design that cannot be rolled back is a design that cannot be changed safely. Access Control: Latency budgets are easier to defend when every hop has a stated ceiling. Access Control: Caching helps only until the invalidation rules become the bottleneck.

Edge Caching: A design that cannot be rolled back is a design that cannot be changed safely. Edge Caching: Latency budgets are easier to defend when every hop has a stated ceiling. Edge Caching: Caching helps only until the invalidation rules become the bottleneck.

Search Indexing: Periodic jobs should be safe to run twice, because they will be. Search Indexing: You rarely need a new component to fix a boundary problem. Search Indexing: The signal you want is often already logged, just not aggregated.

Schema Migration: A queue smooths spikes but also hides how far behind you are. Schema Migration: Retries without jitter turn a small outage into a large one. Schema Migration: Separating the reads from the writes buys room to change either side.

Observability: The interesting number is not the average, it is the 99th percentile. Observability: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Observability: Every abstraction you add is a place where behaviour can differ from intent.

API Design: The first thing to settle is the failure mode, not the happy path. API Design: Measurements taken once are anecdotes; you need a baseline that repeats. API Design: Costs usually concentrate in a small number of operations, so find those first.

Related reading