withColumn
Add or replace one column, and why calling it in a long loop slows Spark down.
On this page
Show code in
Every code block on the page follows this.
You will learn
- How to add a column or replace one in place
- The difference between withColumn, withColumns and withColumnRenamed
- Why a loop of withColumn calls slows Spark down
- How to cast types safely
withColumn(name, expr) returns the DataFrame with one column added, or replaced if the name exists. It is handy for one or two columns; for many, use withColumns or a single select.What it does
It keeps every existing column and appends a new one at the end. If a column with that name already exists, it is replaced in place, keeping its position.
employees.withColumn("tax", F.col("salary") * 0.3)
SELECT *, salary * 0.3 AS tax FROM employees
Step by step
Input
| name | salary |
|---|---|
| Asha | 72000 |
| Ben | 48000 |
Output
| name | salary | band |
|---|---|---|
| Asha | 72000 | high |
| Ben | 48000 | standard |
A new column computed from an existing one. The original DataFrame is unchanged; you get a new one back.
Run the example
from pyspark.sql import functions as F result = (employees .withColumn("band", F.when(F.col("salary") >= 60000, "high").otherwise("standard")) .withColumn("salary", F.col("salary").cast("double")) # replaces in place .withColumnRenamed("dept", "department"))
SELECT id, name, dept AS department, CAST(salary AS DOUBLE) AS salary, manager_id, CASE WHEN salary >= 60000 THEN 'high' ELSE 'standard' END AS band FROM employees
Switch to PySpark to edit and run this example in your browser.
Notice how the SQL version has to list every column to replace one. That convenience is exactly what withColumn gives you.
The withColumn family
| Method | What it does | Since |
|---|---|---|
withColumn(name, col) | Add or replace one column. | 1.3 |
withColumns({name: col, ...}) | Add or replace several columns in one projection. | 3.3 |
withColumnRenamed(old, new) | Rename one column. Does nothing if old does not exist, so typos fail silently. | 1.3 |
withColumnsRenamed({old: new}) | Rename several columns. | 3.4 |
drop("a", "b") | Remove columns. Also silent for names that do not exist. | 1.4 |
The loop problem
Each call wraps the plan in one more Project node. In a loop over hundreds of columns that produces a plan hundreds of levels deep. Spark re-analyses the whole plan on every call, so the cost grows roughly with the square of the number of calls: jobs that should take seconds spend minutes on the driver before a single task starts.
# Slow: one projection per column for c in cols: df = df.withColumn(c, F.trim(F.col(c))) # Fast: one projection for all of them df = df.withColumns({c: F.trim(F.col(c)) for c in cols})
withColumn many times in a loop and recommends select with multiple columns instead.Casting safely
.cast("int") on a string that is not a number returns null in Spark 3, silently. In Spark 4.0, with ANSI mode on by default, the same cast raises an error. If you expect bad values, use try_cast (SQL, or F.expr("try_cast(x AS INT)")), which always returns null on failure, and count the nulls afterwards.
Common mistakes
Expecting withColumn to modify the DataFrame
df.withColumn(...) on its own line does nothing; assign the result.Renaming a column that does not exist
withColumnRenamed and drop ignore unknown names. A typo leaves the old name in place without an error.Generating hundreds of withColumn calls
withColumns or one select.Key takeaways
withColumnadds a column, or replaces it in place if the name exists.- It returns a new DataFrame; always assign the result.
- For many columns use
withColumnsor oneselect. - Renaming or dropping an unknown column is silently ignored.
Check yourself
3 questions1. What does df.withColumn("salary", F.col("salary") * 2) do when salary exists?
Show the answer
Replaces salary in place with the doubled value. withColumn replaces a column with the same name and keeps its position.
2. What happens with df.withColumnRenamed("slary", "pay") if there is no column slary?
Show the answer
Nothing: the DataFrame comes back unchanged. withColumnRenamed is a no-op for names that do not exist, which is why typos go unnoticed.
3. Which is the best way to trim 200 string columns?
Show the answer
df.withColumns({c: F.trim(F.col(c)) for c in cols}). withColumns (or one select) builds a single projection, so planning stays fast.
Practice it
Interview problems that use withColumn: write the PySpark, run it, and get graded on hidden tests.
- Label Late Shipments Easy
- Add a Bonus Column Easy
- Categorize Orders Easy
- Latest Order per Customer Medium
- Running Total by Region Medium
Go deeper
Primary sources: DataFrame.withColumn · DataFrame.withColumns