🔍 Ctrl+K
🐼 Pandas Golden Questionnaire July 2026 • 7 Questions

Pandas Interview Q&A

Covers DataFrame vs Series, GroupBy, merge/join, loc vs iloc, handling nulls, apply functions, and Pandas vs PySpark.

Source: GOLDEN_QUESTIONNAIRE_JULY_2026.pdf • Answers are hidden — click a question to reveal its full interview answer. Use bookmarks + Mark as Complete to track prep.

🟡 Intermediate
Q1

What is a Pandas DataFrame vs a Series? What are the differences?

click to reveal answer
▶

Interview Answer: A Series is a one-dimensional data structure that stores a single column of data with an index, while a DataFrame is a two-dimensional table with rows and multiple columns. A DataFrame is essentially a collection of Series sharing the same index. In my projects, I mainly use DataFrames because ETL pipelines usually involve multiple columns and transformations.

🟡 Intermediate
Q2

How do you read different file formats in Pandas (CSV, JSON, Parquet, Excel)?

click to reveal answer
▶

Interview Answer: Pandas provides dedicated functions to read different file formats. I use read_csv() for CSV files, read_json() for JSON, read_parquet() for Parquet files, and read_excel() for Excel files. Depending on the source system, I choose the appropriate function and then perform data cleaning before further processing.

🟡 Intermediate
Q3

What is the difference between Merge, Join and Concat in Pandas?

click to reveal answer
▶

Interview Answer: Merge combines DataFrames based on common columns, similar to SQL joins. Join is mainly used to combine DataFrames using indexes. Concat simply appends DataFrames either row-wise or column-wise without matching keys. In ETL projects, I mostly use merge to combine customer or transaction datasets.

🟡 Intermediate
Q4

What is the difference between LOC vs ILOC?

click to reveal answer
▶

Interview Answer: loc accesses data using row and column labels, whereas iloc accesses data using integer positions. I use loc when filtering based on column names or indexes, and iloc when selecting rows by position. Both are useful during data validation and exploratory analysis.

🟡 Intermediate
Q5

How do you remove duplicate rows from a DataFrame?

click to reveal answer
▶

Interview Answer: I use the drop_duplicates() function to remove duplicate rows from a DataFrame. If needed, I specify particular columns using the subset parameter to identify duplicates. In ETL pipelines, removing duplicates is an important data quality step before loading data into the target system.

🟡 Intermediate
Q6

How to explode a nested JSON using Pandas?

click to reveal answer
▶

Interview Answer: For nested JSON data, I first flatten the JSON using json_normalize(). If a column contains a list of values, I use the explode() function to convert each list element into a separate row. This makes the data easier to transform and load into relational tables.

🟡 Intermediate
Q7

What is the difference between Pandas vs PySpark?

click to reveal answer
▶

Interview Answer: Pandas processes data on a single machine and is best for small to medium datasets that fit into memory. PySpark is built for distributed computing and can process large-scale data across multiple cluster nodes. In my projects, I use Pandas for local analysis and PySpark for big data ETL pipelines in production.

← All Interview Lessons 🏠 Hub Home