Skip to content
Everyone knows 3 min read · Selection 30 practice problems ↓

withColumn

Add or replace one column, and why calling it in a long loop slows Spark down.

You will learn

  • How to add a column or replace one in place
  • The difference between withColumn, withColumns and withColumnRenamed
  • Why a loop of withColumn calls slows Spark down
  • How to cast types safely

Read first

Comfortable with these? Read on.

TL;DR withColumn(name, expr) returns the DataFrame with one column added, or replaced if the name exists. It is handy for one or two columns; for many, use withColumns or a single select.

What it does

It keeps every existing column and appends a new one at the end. If a column with that name already exists, it is replaced in place, keeping its position.

PySparkSpark SQL
employees.withColumn("tax", F.col("salary") * 0.3)
SELECT *, salary * 0.3 AS tax FROM employees

Step by step

Input

namesalary
Asha72000
Ben48000

Output

namesalaryband
Asha72000high
Ben48000standard

A new column computed from an existing one. The original DataFrame is unchanged; you get a new one back.

Run the example

PySparkSpark SQL
from pyspark.sql import functions as F
result = (employees
    .withColumn("band", F.when(F.col("salary") >= 60000, "high").otherwise("standard"))
    .withColumn("salary", F.col("salary").cast("double"))   # replaces in place
    .withColumnRenamed("dept", "department"))
SELECT id, name, dept AS department,
       CAST(salary AS DOUBLE) AS salary, manager_id,
       CASE WHEN salary >= 60000 THEN 'high' ELSE 'standard' END AS band
FROM employees

Switch to PySpark to edit and run this example in your browser.

Notice how the SQL version has to list every column to replace one. That convenience is exactly what withColumn gives you.

The withColumn family

MethodWhat it doesSince
withColumn(name, col)Add or replace one column.1.3
withColumns({name: col, ...})Add or replace several columns in one projection.3.3
withColumnRenamed(old, new)Rename one column. Does nothing if old does not exist, so typos fail silently.1.3
withColumnsRenamed({old: new})Rename several columns.3.4
drop("a", "b")Remove columns. Also silent for names that do not exist.1.4

The loop problem

Each call wraps the plan in one more Project node. In a loop over hundreds of columns that produces a plan hundreds of levels deep. Spark re-analyses the whole plan on every call, so the cost grows roughly with the square of the number of calls: jobs that should take seconds spend minutes on the driver before a single task starts.

PySpark · Slow, then fast
# Slow: one projection per column
for c in cols:
    df = df.withColumn(c, F.trim(F.col(c)))

# Fast: one projection for all of them
df = df.withColumns({c: F.trim(F.col(c)) for c in cols})
Note: The Spark documentation itself warns against calling withColumn many times in a loop and recommends select with multiple columns instead.

Casting safely

.cast("int") on a string that is not a number returns null in Spark 3, silently. In Spark 4.0, with ANSI mode on by default, the same cast raises an error. If you expect bad values, use try_cast (SQL, or F.expr("try_cast(x AS INT)")), which always returns null on failure, and count the nulls afterwards.

Common mistakes

Expecting withColumn to modify the DataFrame

It returns a new one. df.withColumn(...) on its own line does nothing; assign the result.

Renaming a column that does not exist

withColumnRenamed and drop ignore unknown names. A typo leaves the old name in place without an error.

Generating hundreds of withColumn calls

Use withColumns or one select.

Key takeaways

  • withColumn adds a column, or replaces it in place if the name exists.
  • It returns a new DataFrame; always assign the result.
  • For many columns use withColumns or one select.
  • Renaming or dropping an unknown column is silently ignored.

Check yourself

3 questions

1. What does df.withColumn("salary", F.col("salary") * 2) do when salary exists?

Show the answer

Replaces salary in place with the doubled value. withColumn replaces a column with the same name and keeps its position.

2. What happens with df.withColumnRenamed("slary", "pay") if there is no column slary?

Show the answer

Nothing: the DataFrame comes back unchanged. withColumnRenamed is a no-op for names that do not exist, which is why typos go unnoticed.

3. Which is the best way to trim 200 string columns?

Show the answer

df.withColumns({c: F.trim(F.col(c)) for c in cols}). withColumns (or one select) builds a single projection, so planning stays fast.

Practice it

Interview problems that use withColumn: write the PySpark, run it, and get graded on hidden tests.

Solve: Label Late Shipments →
See all 30 problems →

Go deeper

Primary sources: DataFrame.withColumn · DataFrame.withColumns