Skip to content

33. Learning Path per Student

Difficulty: Medium · Topics: Arrays, Aggregations, Strings, Filtering & Selection

enrollments has student, course and enrolled_on ('YYYY-MM-DD'). Students can retake a course.

For each student return:

Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.

Sample data

enrollments

studentcourseenrolled_on
ashaSpark2025-03-10
ashaSQL2025-01-05
ashaDelta2025-06-01
benPython2025-02-01
benSQL2025-02-01

Expected output

studentpathdistinct_courses
ashaSQL > Spark > Delta3
benPython > SQL2

Hints

Hint 1collect_list does not guarantee any order, even if you sort the DataFrame first: the groupBy shuffle reorders rows.
Hint 2Collect values that sort correctly on their own, such as "2025-03-01|SQL", then F.sort_array them.
Hint 3Join the sorted array into one string with F.array_join, and strip the date prefixes with F.regexp_replace. F.size(F.collect_set("course")) counts distinct courses.

Learn the concepts

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · Learn · All problems