Batch Processing: Spark
When Spark pays off, application anatomy and lazy evaluation, partitions and shuffle, Catalyst and reading query plans, data skew, join strategies and adaptive execution, executor memory, the cost of Python UDFs, writes and small files, caching, Spark Connect