正在加载内容...

963963 Chat Cloud Portal Independent coverage of news

Observability in Practice: Lessons From Real Deployments

By Laura Bennett · · 1267 words
Observability in Practice: Lessons From Real Deployments

Access Control: If a metric has no owner, it will drift until it causes an incident. Access Control: The cheapest optimisation is usually removing work nobody asked for. Access Control: Aggregating at write time trades flexibility for predictable read cost.

Consider observability specifically. The interesting number is not the average, it is the 99th percentile. Observability: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to observability as well.

Serving static bytes is the cheapest thing you can do at the edge. That applies to cost controls as well. In practice, cost controls behaves differently: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. The same reasoning holds for cost controls.

For search indexing, the constraint matters more than the feature list. Configurations should be reviewable in a diff, not only in a console. Teams working on search indexing usually discover this the hard way. The best time to add an index is before the table gets large. Failures are usually correlated, so plan for the shared dependency. This is most visible in search indexing.

Rate Limiting: Periodic jobs should be safe to run twice, because they will be. Rate Limiting: You rarely need a new component to fix a boundary problem. Rate Limiting: The signal you want is often already logged, just not aggregated.

Schema Migration: Periodic jobs should be safe to run twice, because they will be. Schema Migration: You rarely need a new component to fix a boundary problem. Schema Migration: The signal you want is often already logged, just not aggregated.

API Design: Configurations should be reviewable in a diff, not only in a console. API Design: The best time to add an index is before the table gets large. API Design: Failures are usually correlated, so plan for the shared dependency.

Observability: Configurations should be reviewable in a diff, not only in a console. Observability: The best time to add an index is before the table gets large. Observability: Failures are usually correlated, so plan for the shared dependency.

Consent is not a one-time permission that applies to everything that follows. Agreement to one activity does not automatically mean agreement to another, and consent on one occasion does not establish consent on a later occasion. People can set limits, ask to pause or change their minds at any point. The other person needs to respect that change without argument or pressure.

Content Delivery: A design that cannot be rolled back is a design that cannot be changed safely. Content Delivery: Latency budgets are easier to defend when every hop has a stated ceiling. Content Delivery: Caching helps only until the invalidation rules become the bottleneck.

Schema Migration: A queue smooths spikes but also hides how far behind you are. Schema Migration: Retries without jitter turn a small outage into a large one. Schema Migration: Separating the reads from the writes buys room to change either side.

Consider data pipelines specifically. You can often replace a coordination problem with an idempotency key. Data Pipelines: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. That applies to data pipelines as well.

Consider edge caching specifically. Serving static bytes is the cheapest thing you can do at the edge. Edge Caching: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. That applies to edge caching as well.

Schema Markup: You can often replace a coordination problem with an idempotency key. Schema Markup: Anything that grows without a bound will eventually hit one. Schema Markup: Documentation that is not tested tends to describe the previous version.

Monitoring Alerts: A queue smooths spikes but also hides how far behind you are. Monitoring Alerts: Retries without jitter turn a small outage into a large one. Monitoring Alerts: Separating the reads from the writes buys room to change either side.

For release process, the constraint matters more than the feature list. Configurations should be reviewable in a diff, not only in a console. Teams working on release process usually discover this the hard way. The best time to add an index is before the table gets large. Failures are usually correlated, so plan for the shared dependency. This is most visible in release process.

A yes is meaningful when a person can choose freely. Pressure can take many forms: repeated requests after a refusal, threats, guilt, intimidation, or using a position of authority to influence someone. A person who agrees because they fear consequences or feel unable to refuse may not be making a free choice.

Switch the product off and disconnect it from its charger before cleaning. If it uses replaceable batteries, remove them only if the manual directs you to do so. Wash your hands first, then remove visible residue with a clean, soft cloth or rinse the product only when its instructions permit rinsing. Keep water away from seams, buttons, charging contacts and battery compartments unless the maker explicitly says those areas can be exposed.

You can often replace a coordination problem with an idempotency key. That applies to monitoring alerts as well. In practice, monitoring alerts behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for monitoring alerts.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to backup strategy as well. In practice, backup strategy behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for backup strategy.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on load balancing usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

Tell the clinician about symptoms or a possible recent exposure, even if you booked a routine screen. Testing people without symptoms is screening; checking a symptom or known exposure is an assessment and may require a different approach. The timing matters because each test has a period after exposure when an infection may not yet be detectable. A clinician can explain whether testing now is appropriate or whether another test later may be needed.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for edge caching. For edge caching, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on edge caching usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

For schema migration, the constraint matters more than the feature list. Periodic jobs should be safe to run twice, because they will be. Teams working on schema migration usually discover this the hard way. You rarely need a new component to fix a boundary problem. The signal you want is often already logged, just not aggregated. This is most visible in schema migration.

Related reading