Skip to content
Everyone knows 3 min read

Transformations vs actions

Why nothing runs until you call an action, and what lazy evaluation buys Spark.

You will learn

  • The difference between a transformation and an action
  • What lazy evaluation is and why Spark uses it
  • Why the same DataFrame can be computed twice, and how to avoid it
  • How laziness affects where errors appear

Read first

Comfortable with these? Read on.

TL;DR Transformations (select, filter, join, ...) only describe a new DataFrame. Actions (count, collect, write, show) make Spark actually run. Waiting until an action lets Spark optimise the whole chain at once.

Two kinds of operations

Transformations (lazy)Actions (run the job)
select, withColumn, filtercount(), collect(), take(n)
groupBy().agg(), join, orderByshow(), toPandas()
union, distinct, repartitionwrite...save(), saveAsTable
cache() (marks only)foreach, foreachPartition

A transformation returns a new DataFrame immediately, in microseconds, whatever the data size, because it only adds a step to a logical plan. An action returns a value or writes output, so Spark has to compute.

Walk through a pipeline

Line 1 · read

Reading a Parquet table only records the source and reads its schema from file footers or the catalog. No rows are read.

The command

orders = spark.read.parquet("s3://shop/orders")

Line 2 · filter

Adds a Filter node to the plan. Still no data read.

The command

paid = orders.filter(F.col("status") == "PAID")

Line 3 · aggregate

Adds an Aggregate node. The plan is now read, filter, aggregate.

The command

daily = paid.groupBy("order_date").agg(F.sum("amount").alias("revenue"))

Line 4 · action

Now Spark optimises the whole plan: it pushes the status filter into the Parquet scan and reads only the three columns used. Then it runs the job.

The command

daily.write.mode("overwrite").parquet("s3://shop/daily_revenue")

Why lazy?

Seeing the whole pipeline before running it is what makes Catalyst's optimisations possible: filters move down next to the scan, unused columns are never read, consecutive projections merge into one, and joins can pick a strategy knowing what comes after. Running each line eagerly, as pandas does, would read everything first.

The recomputation trap

A DataFrame is a recipe, not a result. Every action re-runs the recipe from the source:

PySpark · Two actions, two full reads
clean = raw.filter(...).withColumn(...)   # expensive
clean.count()                               # job 1: reads and cleans
clean.write.parquet("out")                  # job 2: reads and cleans again

If a DataFrame is used by several actions, either cache it (see the Caching lesson) or restructure so a single action does the work. A common unnecessary action is a count() just to log a number.

Where errors show up

  • Analysis errors (a missing column, a type mismatch) appear immediately when you write the transformation, because Spark resolves names eagerly.
  • Runtime errors (a bad cast in ANSI mode, a division by zero, a corrupt file, an out-of-memory) appear only at the action, often many lines later. The stack trace points at the action, not at the transformation that caused it.
Tip: when debugging, add a temporary limit(10).show() after a suspicious step to force evaluation there. Remove it afterwards.

Common mistakes

Timing transformations

Measuring how long a filter "takes" measures nothing; time the action.

Calling several actions on an uncached DataFrame

Each action recomputes everything from the source.

Debug count() calls left in production

Each one is a full extra job.

Key takeaways

  • Transformations build a plan; actions run it.
  • Laziness lets Catalyst optimise the whole pipeline.
  • Each action recomputes from the source unless you cache.
  • Runtime errors surface at the action, not where they were caused.

Check yourself

3 questions

1. Which is an action?

Show the answer

count(). count returns a value to the driver, so Spark must compute.

2. A DataFrame is written and then counted, with no cache. How many times is the source read?

Show the answer

Twice. Each action re-runs the plan from the source.

3. When does a missing column error appear?

Show the answer

Immediately, when the transformation is defined. Spark analyses names eagerly, so unresolved columns fail right away.

Go deeper

Primary sources: RDD programming guide: lazy evaluation · Spark SQL guide