Frontend delivery, platform operations for the cluster, monitoring with Prometheus and Grafana, and AI/ML infrastructure: the layer between service code and cloud primitives
Frontend Delivery
Build and bundle budgets, asset versioning and cache headers, atomic deploys, RUM and Core Web Vitals, browser error monitoring, locales on CDN, progressive rollout, delivery security
Platform & Operations
Cluster networking and DNS, network policies, scheduler and pod placement, container runtime security, quotas and multi-tenancy, production debugging, jobs and schedules, backups, cluster upgrades
AI & ML Infrastructure
GPUs in the cluster and job scheduling, distributed training and checkpoints, dataset and weight storage, cost and utilization, model cold starts, response caching, data isolation, AI feature degradation
Prometheus & Metrics
The pull model and data model, metric types and PromQL, service discovery and relabeling, exporters and instrumentation, histograms and quantiles, TSDB internals, cardinality, rules and Alertmanager, long-term storage and running a highly available pair
Grafana & Dashboards
Data sources and proxy access, panels and visualizations, dashboard anatomy and variables, repeats and transformations, units and thresholds, $__rate_interval, Grafana Alerting, provisioning and access control, dashboard performance, exemplars and dynamic dashboards