Skip to content

Data engineering & AI · Corporate training

Apache Spark & PySpark

Process larger datasets with distributed transformations, Spark SQL, and streaming foundations.

  • Intermediate
  • Live online and in person
  • Duration tailored to your team
Discuss this trainingExplore sample curriculum ↓

Sample curriculum

What your team can learn.

This outline is a starting point. Modules, exercises, and depth are adapted to your team's experience and project requirements.

01Distributed processing
  • Spark architecture and execution
  • Sessions, dataframes, and schemas
  • Reading and writing common data formats
02Transformations & SQL
  • Filters, expressions, and aggregations
  • Joins and window functions
  • Spark SQL and reusable transformations
03Performance & reliability
  • Lazy evaluation and execution plans
  • Partitions, shuffles, and caching
  • Data quality and failure investigation
04Streaming & pipeline design
  • Structured Streaming concepts
  • Checkpoints and incremental processing
  • Project review and performance comparison

Hands-on capstone

Put the learning into practice.

Build a batch processing pipeline and inspect its execution plan, partitions, and output quality.

How the training works

Live demonstrations, guided exercises, pair work, and project reviews connect the concepts to practical implementation.

We agree on your learning goals, prerequisites, group size, delivery format, and schedule before shaping the final curriculum and proposal.

Your trainer

I'm Ragav Kumar V, a corporate trainer and mentor with 10+ years of training experience since 2016, 100+ batches delivered, and 3,000+ professionals trained across full-stack development, data engineering, and Cloud & DevOps.

Continue your learning path.