正在加载内容...

963963 Chat News Review Independent coverage of news

Seven Things to Check Before Choosing Data Pipelines

By David Kim · · 1263 words
Seven Things to Check Before Choosing Data Pipelines

Release Process: If a metric has no owner, it will drift until it causes an incident. Release Process: The cheapest optimisation is usually removing work nobody asked for. Release Process: Aggregating at write time trades flexibility for predictable read cost.

Queue Design: The interesting number is not the average, it is the 99th percentile. Queue Design: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Queue Design: Every abstraction you add is a place where behaviour can differ from intent.

The interesting number is not the average, it is the 99th percentile. That applies to api design as well. In practice, api design behaves differently: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. The same reasoning holds for api design.

Data Pipelines: A queue smooths spikes but also hides how far behind you are. Data Pipelines: Retries without jitter turn a small outage into a large one. Data Pipelines: Separating the reads from the writes buys room to change either side.

Search Indexing: A queue smooths spikes but also hides how far behind you are. Search Indexing: Retries without jitter turn a small outage into a large one. Search Indexing: Separating the reads from the writes buys room to change either side.

Boundaries can change with circumstances, health, trust or preference. Partners can check in before a new activity or after an experience, without treating a previous agreement as permanent. Digital boundaries deserve the same care as in-person ones: discuss private messages, location sharing, passwords and images. Consent to receive or make an image is not permission to forward it.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to cost controls as well. In practice, cost controls behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for cost controls.

Separate a boundary from a preference where you can. A preference describes something you like or would choose; a boundary describes what you are not willing to do, or what you need in order to feel comfortable. Both are useful information, but a boundary should not be treated as an opening offer to negotiate. You can say, “I’m not comfortable with that,” without supplying a detailed reason.

Consider schema migration specifically. A design that cannot be rolled back is a design that cannot be changed safely. Schema Migration: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. That applies to schema migration as well.

Release Process: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. That applies to release process as well. In practice, release process behaves differently: Aggregating at write time trades flexibility for predictable read cost.

For access control, the constraint matters more than the feature list. The first thing to settle is the failure mode, not the happy path. Teams working on access control usually discover this the hard way. Measurements taken once are anecdotes; you need a baseline that repeats. Costs usually concentrate in a small number of operations, so find those first. This is most visible in access control.

Search Indexing: Configurations should be reviewable in a diff, not only in a console. Search Indexing: The best time to add an index is before the table gets large. Search Indexing: Failures are usually correlated, so plan for the shared dependency.

Crawl Budget: A design that cannot be rolled back is a design that cannot be changed safely. Crawl Budget: Latency budgets are easier to defend when every hop has a stated ceiling. Crawl Budget: Caching helps only until the invalidation rules become the bottleneck.

It can help to prepare a short sentence and a next step. For instance: “I want to take things slowly, so let’s check in before anything changes,” or “I don’t want photos taken or shared.” If you are unsure what you want, say so. “I’m still working that out, and I want to pause for now” communicates a limit without requiring you to settle every future question.

Data Pipelines: The first thing to settle is the failure mode, not the happy path. Data Pipelines: Measurements taken once are anecdotes; you need a baseline that repeats. Data Pipelines: Costs usually concentrate in a small number of operations, so find those first.

Crawl Budget: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. That applies to crawl budget as well. In practice, crawl budget behaves differently: Separating the reads from the writes buys room to change either side.

Serving static bytes is the cheapest thing you can do at the edge. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. A schema is an interface; changing it is a migration, not an edit. Teams working on queue design usually discover this the hard way. Track the denominator as carefully as the numerator.

Access Control: A design that cannot be rolled back is a design that cannot be changed safely. Access Control: Latency budgets are easier to defend when every hop has a stated ceiling. Access Control: Caching helps only until the invalidation rules become the bottleneck.

Observability: A queue smooths spikes but also hides how far behind you are. Observability: Retries without jitter turn a small outage into a large one. Observability: Separating the reads from the writes buys room to change either side.

Log Analysis: Serving static bytes is the cheapest thing you can do at the edge. Log Analysis: A schema is an interface; changing it is a migration, not an edit. Log Analysis: Track the denominator as carefully as the numerator.

Communication does not have to follow a script. Partners can discuss boundaries and expectations before an intimate situation, then check in again if circumstances or preferences change. Nonverbal communication can provide context, but gestures or body language may be misread; they should not be treated as a substitute for clear agreement when there is doubt. People who communicate in different ways can agree on accessible ways to express yes, no and pause.

Tell the clinician about symptoms or a possible recent exposure, even if you booked a routine screen. Testing people without symptoms is screening; checking a symptom or known exposure is an assessment and may require a different approach. The timing matters because each test has a period after exposure when an infection may not yet be detectable. A clinician can explain whether testing now is appropriate or whether another test later may be needed.

For backup strategy, the constraint matters more than the feature list. If a metric has no owner, it will drift until it causes an incident. Teams working on backup strategy usually discover this the hard way. The cheapest optimisation is usually removing work nobody asked for. Aggregating at write time trades flexibility for predictable read cost. This is most visible in backup strategy.

Cloud Infrastructure: Configurations should be reviewable in a diff, not only in a console. Cloud Infrastructure: The best time to add an index is before the table gets large. Cloud Infrastructure: Failures are usually correlated, so plan for the shared dependency.

Related reading