正在加载内容...

963963 Chat News Review Independent coverage of news

Observability Compared: What Actually Matters

By David Kim · · 1213 words
Observability Compared: What Actually Matters

Access Control: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. That applies to access control as well. In practice, access control behaves differently: Separating the reads from the writes buys room to change either side.

Teams working on search indexing usually discover this the hard way. Serving static bytes is the cheapest thing you can do at the edge. A schema is an interface; changing it is a migration, not an edit. This is most visible in search indexing. Consider search indexing specifically. Track the denominator as carefully as the numerator.

If the rollback plan needs a meeting, it is not a rollback plan. That applies to edge caching as well. In practice, edge caching behaves differently: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. The same reasoning holds for edge caching.

Schema Migration: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. That applies to schema migration as well. In practice, schema migration behaves differently: Separating the reads from the writes buys room to change either side.

Partners may have different preferences. They can discuss whether there is an option both freely want, but neither person owes a compromise involving their body, safety or privacy. If there is no mutually acceptable option, stopping or not doing the activity is a valid outcome. A difference in boundaries can also reveal a broader mismatch in expectations; that does not make either person’s limit less legitimate.

Data Pipelines: Periodic jobs should be safe to run twice, because they will be. Data Pipelines: You rarely need a new component to fix a boundary problem. Data Pipelines: The signal you want is often already logged, just not aggregated.

Conversation about consent can include practical safety decisions, such as boundaries, contraception and protection from sexually transmitted infections. These discussions do not replace medical advice, and agreement about one safety measure does not imply agreement to anything else. If plans or conditions change, revisit the agreement rather than assuming earlier consent still applies.

Consider access control specifically. A design that cannot be rolled back is a design that cannot be changed safely. Access Control: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. That applies to access control as well.

Release Process: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. That applies to release process as well. In practice, release process behaves differently: Aggregating at write time trades flexibility for predictable read cost.

Consider content delivery specifically. The interesting number is not the average, it is the 99th percentile. Content Delivery: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to content delivery as well.

Storage Tiers: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. That applies to storage tiers as well. In practice, storage tiers behaves differently: Failures are usually correlated, so plan for the shared dependency.

Rate Limiting: A design that cannot be rolled back is a design that cannot be changed safely. Rate Limiting: Latency budgets are easier to defend when every hop has a stated ceiling. Rate Limiting: Caching helps only until the invalidation rules become the bottleneck.

Data Pipelines: A design that cannot be rolled back is a design that cannot be changed safely. Data Pipelines: Latency budgets are easier to defend when every hop has a stated ceiling. Data Pipelines: Caching helps only until the invalidation rules become the bottleneck.

Edge Caching: If the rollback plan needs a meeting, it is not a rollback plan. Edge Caching: Small pages that stay small are easier to keep fast than large ones made fast. Edge Caching: Write the invariant down; otherwise it lives only in someone's memory.

Observability: You can often replace a coordination problem with an idempotency key. Observability: Anything that grows without a bound will eventually hit one. Observability: Documentation that is not tested tends to describe the previous version.

Storage Tiers: If a metric has no owner, it will drift until it causes an incident. Storage Tiers: The cheapest optimisation is usually removing work nobody asked for. Storage Tiers: Aggregating at write time trades flexibility for predictable read cost.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to data pipelines as well. In practice, data pipelines behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for data pipelines.

If the rollback plan needs a meeting, it is not a rollback plan. That applies to access control as well. In practice, access control behaves differently: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. The same reasoning holds for access control.

Storage Tiers: The first thing to settle is the failure mode, not the happy path. Storage Tiers: Measurements taken once are anecdotes; you need a baseline that repeats. Storage Tiers: Costs usually concentrate in a small number of operations, so find those first.

The first thing to settle is the failure mode, not the happy path. This is most visible in backup strategy. Consider backup strategy specifically. Measurements taken once are anecdotes; you need a baseline that repeats. Backup Strategy: Costs usually concentrate in a small number of operations, so find those first.

For edge caching, the constraint matters more than the feature list. Periodic jobs should be safe to run twice, because they will be. Teams working on edge caching usually discover this the hard way. You rarely need a new component to fix a boundary problem. The signal you want is often already logged, just not aggregated. This is most visible in edge caching.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for access control. For access control, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on access control usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

Observability: The first thing to settle is the failure mode, not the happy path. Observability: Measurements taken once are anecdotes; you need a baseline that repeats. Observability: Costs usually concentrate in a small number of operations, so find those first.

Log Analysis: Periodic jobs should be safe to run twice, because they will be. Log Analysis: You rarely need a new component to fix a boundary problem. Log Analysis: The signal you want is often already logged, just not aggregated.

Related reading