About CodeForData
CodeForData is a practice site for the PySpark questions asked in data engineering interviews. Each problem gives you one or more DataFrames, describes the result you need, and checks your code against a sample and against hidden test cases built around edge cases: nulls, ties, empty groups and duplicate keys.
How it works
- Run executes your code instantly in a PySpark DataFrame emulator inside your browser, against the sample data shown on the page. There is nothing to install and no cluster to wait for.
- Submit grades your code on our server against every test case, including the hidden ones. Once you pass, the reference solution unlocks with an explanation of why it works.
- You can practise without an account. Logging in keeps your solved problems, code and submissions in sync across devices.
What you can practise
Filtering and selection, conditional logic, null handling, aggregations, strings, dates, joins, arrays, window functions, and pivot, unpivot and rollup, from easy warm-ups to senior-level patterns such as sessionization, gaps and islands and SCD Type 2. Browse all problems or follow the PySpark interview roadmap.
The emulator
The emulator implements a subset of the PySpark DataFrame API (the everyday functions interviews focus on). It is not Apache Spark, so performance questions about partitions and shuffles are outside its scope. PySpark is a trademark of the Apache Software Foundation; CodeForData is not affiliated with it.
Who runs it
CodeForData is an independent project operated by Suraj Anand in India. Questions, corrections and feedback are welcome on the Contact page.