正在加载内容...

963963 Chat Cloud Portal Independent coverage of news

Data Pipelines Explained Without the Jargon

By James Whitfield · · 1199 words
Data Pipelines Explained Without the Jargon

Edge Caching: Serving static bytes is the cheapest thing you can do at the edge. Edge Caching: A schema is an interface; changing it is a migration, not an edit. Edge Caching: Track the denominator as carefully as the numerator.

The interesting number is not the average, it is the 99th percentile. That applies to rate limiting as well. In practice, rate limiting behaves differently: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. The same reasoning holds for rate limiting.

Queue Design: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to queue design as well. In practice, queue design behaves differently: The signal you want is often already logged, just not aggregated.

Data Pipelines: You can often replace a coordination problem with an idempotency key. Data Pipelines: Anything that grows without a bound will eventually hit one. Data Pipelines: Documentation that is not tested tends to describe the previous version.

Content Delivery: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to content delivery as well. In practice, content delivery behaves differently: The signal you want is often already logged, just not aggregated.

Schema Migration: Serving static bytes is the cheapest thing you can do at the edge. Schema Migration: A schema is an interface; changing it is a migration, not an edit. Schema Migration: Track the denominator as carefully as the numerator.

Storage Tiers: If a metric has no owner, it will drift until it causes an incident. Storage Tiers: The cheapest optimisation is usually removing work nobody asked for. Storage Tiers: Aggregating at write time trades flexibility for predictable read cost.

A queue smooths spikes but also hides how far behind you are. This is most visible in rate limiting. Consider rate limiting specifically. Retries without jitter turn a small outage into a large one. Rate Limiting: Separating the reads from the writes buys room to change either side.

In practice, load balancing behaves differently: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Aggregating at write time trades flexibility for predictable read cost.

Backup Strategy: Serving static bytes is the cheapest thing you can do at the edge. Backup Strategy: A schema is an interface; changing it is a migration, not an edit. Backup Strategy: Track the denominator as carefully as the numerator.

If a metric has no owner, it will drift until it causes an incident. This is most visible in queue design. Consider queue design specifically. The cheapest optimisation is usually removing work nobody asked for. Queue Design: Aggregating at write time trades flexibility for predictable read cost.

Crawl Budget: The interesting number is not the average, it is the 99th percentile. Crawl Budget: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Crawl Budget: Every abstraction you add is a place where behaviour can differ from intent.

Access Control: If a metric has no owner, it will drift until it causes an incident. Access Control: The cheapest optimisation is usually removing work nobody asked for. Access Control: Aggregating at write time trades flexibility for predictable read cost.

If the rollback plan needs a meeting, it is not a rollback plan. The same reasoning holds for storage tiers. For storage tiers, the constraint matters more than the feature list. Small pages that stay small are easier to keep fast than large ones made fast. Teams working on storage tiers usually discover this the hard way. Write the invariant down; otherwise it lives only in someone's memory.

The interesting number is not the average, it is the 99th percentile. That applies to search indexing as well. In practice, search indexing behaves differently: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. The same reasoning holds for search indexing.

Rate Limiting: The first thing to settle is the failure mode, not the happy path. Rate Limiting: Measurements taken once are anecdotes; you need a baseline that repeats. Rate Limiting: Costs usually concentrate in a small number of operations, so find those first.

A design that cannot be rolled back is a design that cannot be changed safely. The same reasoning holds for monitoring alerts. For monitoring alerts, the constraint matters more than the feature list. Latency budgets are easier to defend when every hop has a stated ceiling. Teams working on monitoring alerts usually discover this the hard way. Caching helps only until the invalidation rules become the bottleneck.

Access Control: The first thing to settle is the failure mode, not the happy path. Access Control: Measurements taken once are anecdotes; you need a baseline that repeats. Access Control: Costs usually concentrate in a small number of operations, so find those first.

Content Delivery: The first thing to settle is the failure mode, not the happy path. Content Delivery: Measurements taken once are anecdotes; you need a baseline that repeats. Content Delivery: Costs usually concentrate in a small number of operations, so find those first.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to cost controls as well. In practice, cost controls behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for cost controls.

Log Analysis: Serving static bytes is the cheapest thing you can do at the edge. Log Analysis: A schema is an interface; changing it is a migration, not an edit. Log Analysis: Track the denominator as carefully as the numerator.

Serving static bytes is the cheapest thing you can do at the edge. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. A schema is an interface; changing it is a migration, not an edit. Teams working on queue design usually discover this the hard way. Track the denominator as carefully as the numerator.

Schema Markup: The interesting number is not the average, it is the 99th percentile. Schema Markup: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Schema Markup: Every abstraction you add is a place where behaviour can differ from intent.

Teams working on monitoring alerts usually discover this the hard way. If the rollback plan needs a meeting, it is not a rollback plan. Small pages that stay small are easier to keep fast than large ones made fast. This is most visible in monitoring alerts. Consider monitoring alerts specifically. Write the invariant down; otherwise it lives only in someone's memory.

Related reading