PySpark Interview Q&A
Covers Spark architecture, RDDs, lazy evaluation, transformations vs actions, cache vs persist, partitioning, skew, and salting.
Source: GOLDEN_QUESTIONNAIRE_JULY_2026.pdf • Answers are hidden — click a question to reveal its full interview answer. Use bookmarks + Mark as Complete to track prep.
Explain Spark Architecture in detail.
click to reveal answerInterview Answer: Spark follows a master-worker architecture. When a Spark job starts, the Driver creates the execution plan and coordinates the job. The Cluster Manager allocates resources, and Executors on worker nodes execute the tasks and store data in memory. This distributed architecture enables Spark to process large datasets efficiently.
What happens after a spark-submit?
click to reveal answerInterview Answer: After spark-submit, the Driver program starts and requests resources from the Cluster Manager. Executors are launched on worker nodes, and the Driver builds a DAG (Directed Acyclic Graph) based on the transformations. When an action is triggered, Spark schedules tasks and executors process the data in parallel before returning the results.
RDD vs DataFrame vs Dataset (4 differences with examples)
click to reveal answerInterview Answer: RDD is the low-level distributed collection with full control but no optimization. DataFrame is a structured table with schema and uses the Catalyst Optimizer for better performance. Dataset combines DataFrame optimization with compile-time type safety, mainly in Scala and Java. In PySpark, I mostly use DataFrames because they are faster and easier to work with.
What is Lazy Evaluation in Spark? Why is it important?
click to reveal answerInterview Answer: Spark doesn't execute transformations immediately. Instead, it records them in a DAG and executes only when an action like show(), count(), or collect() is called. This is called Lazy Evaluation. It helps Spark optimize execution, reduce unnecessary computations, and improve overall performance.
Narrow Transformation vs Wide Transformation (4 differences with examples)
click to reveal answerInterview Answer: Narrow transformations process data within the same partition and don't require data movement, such as map(), filter(), and flatMap(). Wide transformations require shuffling data across partitions, such as groupBy(), join(), and reduceByKey(). Narrow transformations are faster because they avoid network communication.
What is a Sort-Merge Join vs Shuffle-Hash Join?
click to reveal answerInterview Answer: Sort-Merge Join first shuffles and sorts both datasets before joining them. It's best for large datasets and is Spark's default join strategy. Shuffle-Hash Join also shuffles data but builds a hash table on the smaller side instead of sorting. It is faster when one dataset is significantly smaller.
Cache vs Persist (4 differences and when to use each)
click to reveal answerInterview Answer: Both cache() and persist() store data for reuse. cache() stores data only in memory using the default storage level. persist() allows different storage options such as memory, disk, or both. I use cache() for frequently accessed small datasets and persist() when the dataset is too large to fit entirely in memory.
Repartition vs Coalesce (4 differences and when to use each)
click to reveal answerInterview Answer: Repartition() increases or decreases partitions and performs a full shuffle, making it suitable for balancing data before joins. Coalesce() mainly reduces partitions without a full shuffle, making it faster. I use repartition() for better parallelism and coalesce() before writing output files to reduce the number of files.
What is Data Skewness? What causes it and how do you handle it?
click to reveal answerInterview Answer: Data skewness occurs when a few partitions contain much more data than others, causing some executors to take longer. It is usually caused by uneven key distribution during joins or aggregations. I handle it using techniques like salting, broadcast joins, repartitioning, or filtering skewed data.
What is Salting? Why and how do you implement it in Spark? Provide code example.
click to reveal answerInterview Answer: Salting is a technique used to handle data skew by adding a random value to skewed keys, distributing records across multiple partitions. This prevents one executor from processing all records for the same key and improves join performance.
from pyspark.sql.functions import concat, lit, rand, floor
df1 = df1.withColumn(
"salted_key",
concat(df1.id, lit("_"), floor(rand()*5))
)
What are the optimization techniques to improve Spark performance?
click to reveal answerInterview Answer: I improve Spark performance by using partitioning, broadcast joins, caching, avoiding unnecessary shuffle operations, selecting only required columns, and using efficient file formats like Parquet. I also tune executor memory, optimize partition sizes, and monitor jobs using the Spark UI.
What is Broadcast Join vs Salting?
click to reveal answerInterview Answer: Broadcast Join is used when one table is small. Spark sends the small table to all executors, avoiding expensive shuffles. Salting is used to solve data skew by distributing heavily repeated keys across multiple partitions. I use broadcast for small lookup tables and salting when joins are slowed down by skewed data.