Skip to content
Everyone knows 3 min read · Selection 24 practice problems ↓

orderBy

Sort rows ascending or descending, control where nulls go, and know what a global sort costs.

You will learn

  • How to sort by several columns and directions
  • How to control where nulls go
  • Why a global sort is expensive, and when sortWithinPartitions is enough
  • Why limit after orderBy is cheap

Read first

Comfortable with these? Read on.

TL;DR orderBy (alias sort) sorts the whole DataFrame. Ascending puts nulls first and descending puts them last, unless you say otherwise. A global sort needs a shuffle, so use it only when you need total order.

What it does

Sorts rows by one or more columns. Each column can be ascending or descending, and you can choose whether nulls come first or last.

Step by step

Input

namedeptsalary
BenSales48000
Faridnull39000
AshaEngineering72000
EshaSales50000

Output

namedeptsalary
AshaEngineering72000
EshaSales50000
BenSales48000
Faridnull39000

Sorted by dept ascending with nulls last, then salary descending within each department.

Run the example

PySparkSpark SQL
from pyspark.sql import functions as F
result = employees.orderBy(
    F.col("dept").asc_nulls_last(),
    F.col("salary").desc(),
)
SELECT *
FROM employees
ORDER BY dept ASC NULLS LAST, salary DESC

Switch to PySpark to edit and run this example in your browser.

Where nulls go

DirectionDefaultOverride
AscendingNulls firstasc_nulls_last() / ASC NULLS LAST
DescendingNulls lastdesc_nulls_first() / DESC NULLS FIRST

Other databases differ: PostgreSQL treats nulls as larger than any value, so its default is the opposite. Say it explicitly when results must match another system.

What a global sort costs

To sort the whole DataFrame, Spark first samples the data to pick range boundaries, then shuffles every row into the partition for its range, then sorts each partition. That is a full shuffle plus a sort, and extra sampling jobs you will see in the Spark UI.

orderBy / sort

Total order across all partitions. Needs a range-partitioning shuffle.

sortWithinPartitions

Sorts each partition independently. No shuffle. Enough before writing files, so each file is sorted for better min/max skipping.
Tip: orderBy(...).limit(10) is cheap: Spark plans it as TakeOrderedAndProject, keeping only the top 10 per partition before merging, instead of sorting everything.

Ties and stability

Rows that tie on every sort key can come back in any order, and the order can change between runs. If a result must be deterministic, for example for pagination or tests, add a unique column as the last sort key.

Common mistakes

Sorting before a groupBy or join

The next shuffle destroys the order. Sort last.

Assuming the sort survives a write

Files written from a sorted DataFrame are not read back in a guaranteed order. Sort in the query that needs it.

Leaving ties unresolved

Add a unique tiebreaker when order must be stable.

Key takeaways

  • orderBy sorts globally; ascending puts nulls first by default.
  • Use asc_nulls_last / NULLS LAST to control null placement.
  • A global sort shuffles everything; sortWithinPartitions does not.
  • Sort plus limit runs as a cheap top-K.

Check yourself

3 questions

1. With default settings, where do nulls go in orderBy(F.col("x").desc())?

Show the answer

Last. Spark puts nulls first for ascending order and last for descending order by default.

2. Which needs a shuffle?

Show the answer

orderBy. A global orderBy range-partitions the data. sortWithinPartitions sorts each partition where it is.

3. Two rows tie on every sort key. What is guaranteed about their order?

Show the answer

Nothing. Spark makes no guarantee for ties. Add a unique tiebreaker for deterministic output.

Practice it

Interview problems that use orderBy: write the PySpark, run it, and get graded on hidden tests.

Solve: Sort Delivery Options →
See all 24 problems →

Go deeper

Primary sources: DataFrame.orderBy · ORDER BY clause