Relational databases in depth, NoSQL and caches, analytical and search engines, message brokers and streaming, batch processing with Spark, working with data from the application: pools, ORM, migrations
Relational Databases
Indexes and query plans, statistics, MVCC and vacuum, isolation levels in practice, locks and deadlocks, partitioning, replication and replica lag, sharding
NoSQL & Caches
Redis and its data structures, eviction and persistence, caching strategies and stampedes, document and wide-column stores, consistency models
Analytics & Search
Columnar stores and ClickHouse, OLAP versus OLTP, data formats, full-text search and Elasticsearch, vector indexes, object storage
Brokers & Streaming
Kafka in depth: partitions, offsets, delivery guarantees; RabbitMQ, NATS and SQS; idempotency and deduplication, DLQ, schema registry, CDC and outbox, stream processing
Data Access from the App
Connection pools and PgBouncer, N+1 and ORM, zero-downtime migrations, cache invalidation, service-level transactions, testing against a database
Batch Processing: Spark
When Spark pays off, application anatomy and lazy evaluation, partitions and shuffle, Catalyst and reading query plans, data skew, join strategies and adaptive execution, executor memory, the cost of Python UDFs, writes and small files, caching, Spark Connect