🔍 Ctrl+K
⚡ PySpark Golden Questionnaire July 2026 • 12 Questions

PySpark Interview Q&A

Covers Spark architecture, RDDs, lazy evaluation, transformations vs actions, cache vs persist, partitioning, skew, and salting.

Source: GOLDEN_QUESTIONNAIRE_JULY_2026.pdf • Answers are hidden — click a question to reveal its full interview answer. Use bookmarks + Mark as Complete to track prep.

🟡 Intermediate
Q1

Explain Spark Architecture in detail.

click to reveal answer
▶

Interview Answer: Spark follows a master-worker architecture. When a Spark job starts, the Driver creates the execution plan and coordinates the job. The Cluster Manager allocates resources, and Executors on worker nodes execute the tasks and store data in memory. This distributed architecture enables Spark to process large datasets efficiently.

🟡 Intermediate
Q2

What happens after a spark-submit?

click to reveal answer
▶

Interview Answer: After spark-submit, the Driver program starts and requests resources from the Cluster Manager. Executors are launched on worker nodes, and the Driver builds a DAG (Directed Acyclic Graph) based on the transformations. When an action is triggered, Spark schedules tasks and executors process the data in parallel before returning the results.

🟡 Intermediate
Q3

RDD vs DataFrame vs Dataset (4 differences with examples)

click to reveal answer
▶

Interview Answer: RDD is the low-level distributed collection with full control but no optimization. DataFrame is a structured table with schema and uses the Catalyst Optimizer for better performance. Dataset combines DataFrame optimization with compile-time type safety, mainly in Scala and Java. In PySpark, I mostly use DataFrames because they are faster and easier to work with.

🟡 Intermediate
Q4

What is Lazy Evaluation in Spark? Why is it important?

click to reveal answer
▶

Interview Answer: Spark doesn't execute transformations immediately. Instead, it records them in a DAG and executes only when an action like show(), count(), or collect() is called. This is called Lazy Evaluation. It helps Spark optimize execution, reduce unnecessary computations, and improve overall performance.

🟡 Intermediate
Q5

Narrow Transformation vs Wide Transformation (4 differences with examples)

click to reveal answer
▶

Interview Answer: Narrow transformations process data within the same partition and don't require data movement, such as map(), filter(), and flatMap(). Wide transformations require shuffling data across partitions, such as groupBy(), join(), and reduceByKey(). Narrow transformations are faster because they avoid network communication.

🟡 Intermediate
Q6

What is a Sort-Merge Join vs Shuffle-Hash Join?

click to reveal answer
▶

Interview Answer: Sort-Merge Join first shuffles and sorts both datasets before joining them. It's best for large datasets and is Spark's default join strategy. Shuffle-Hash Join also shuffles data but builds a hash table on the smaller side instead of sorting. It is faster when one dataset is significantly smaller.

🟡 Intermediate
Q7

Cache vs Persist (4 differences and when to use each)

click to reveal answer
▶

Interview Answer: Both cache() and persist() store data for reuse. cache() stores data only in memory using the default storage level. persist() allows different storage options such as memory, disk, or both. I use cache() for frequently accessed small datasets and persist() when the dataset is too large to fit entirely in memory.

🟡 Intermediate
Q8

Repartition vs Coalesce (4 differences and when to use each)

click to reveal answer
▶

Interview Answer: Repartition() increases or decreases partitions and performs a full shuffle, making it suitable for balancing data before joins. Coalesce() mainly reduces partitions without a full shuffle, making it faster. I use repartition() for better parallelism and coalesce() before writing output files to reduce the number of files.

🟡 Intermediate
Q9

What is Data Skewness? What causes it and how do you handle it?

click to reveal answer
▶

Interview Answer: Data skewness occurs when a few partitions contain much more data than others, causing some executors to take longer. It is usually caused by uneven key distribution during joins or aggregations. I handle it using techniques like salting, broadcast joins, repartitioning, or filtering skewed data.

🟡 Intermediate
Q10

What is Salting? Why and how do you implement it in Spark? Provide code example.

click to reveal answer
▶

Interview Answer: Salting is a technique used to handle data skew by adding a random value to skewed keys, distributing records across multiple partitions. This prevents one executor from processing all records for the same key and improves join performance.

Example
from pyspark.sql.functions import concat, lit, rand, floor
df1 = df1.withColumn(
    "salted_key",
    concat(df1.id, lit("_"), floor(rand()*5))
)
🟡 Intermediate
Q11

What are the optimization techniques to improve Spark performance?

click to reveal answer
▶

Interview Answer: I improve Spark performance by using partitioning, broadcast joins, caching, avoiding unnecessary shuffle operations, selecting only required columns, and using efficient file formats like Parquet. I also tune executor memory, optimize partition sizes, and monitor jobs using the Spark UI.

🟡 Intermediate
Q12

What is Broadcast Join vs Salting?

click to reveal answer
▶

Interview Answer: Broadcast Join is used when one table is small. Spark sends the small table to all executors, avoiding expensive shuffles. Salting is used to solve data skew by distributing heavily repeated keys across multiple partitions. I use broadcast for small lookup tables and salting when joins are slowed down by skewed data.

← All Interview Lessons 🏠 Hub Home