Pandas Interview Q&A
Covers DataFrame vs Series, GroupBy, merge/join, loc vs iloc, handling nulls, apply functions, and Pandas vs PySpark.
Source: GOLDEN_QUESTIONNAIRE_JULY_2026.pdf • Answers are hidden — click a question to reveal its full interview answer. Use bookmarks + Mark as Complete to track prep.
What is a Pandas DataFrame vs a Series? What are the differences?
click to reveal answerInterview Answer: A Series is a one-dimensional data structure that stores a single column of data with an index, while a DataFrame is a two-dimensional table with rows and multiple columns. A DataFrame is essentially a collection of Series sharing the same index. In my projects, I mainly use DataFrames because ETL pipelines usually involve multiple columns and transformations.
How do you read different file formats in Pandas (CSV, JSON, Parquet, Excel)?
click to reveal answerInterview Answer: Pandas provides dedicated functions to read different file formats. I use read_csv() for CSV files, read_json() for JSON, read_parquet() for Parquet files, and read_excel() for Excel files. Depending on the source system, I choose the appropriate function and then perform data cleaning before further processing.
What is the difference between Merge, Join and Concat in Pandas?
click to reveal answerInterview Answer: Merge combines DataFrames based on common columns, similar to SQL joins. Join is mainly used to combine DataFrames using indexes. Concat simply appends DataFrames either row-wise or column-wise without matching keys. In ETL projects, I mostly use merge to combine customer or transaction datasets.
What is the difference between LOC vs ILOC?
click to reveal answerInterview Answer: loc accesses data using row and column labels, whereas iloc accesses data using integer positions. I use loc when filtering based on column names or indexes, and iloc when selecting rows by position. Both are useful during data validation and exploratory analysis.
How do you remove duplicate rows from a DataFrame?
click to reveal answerInterview Answer: I use the drop_duplicates() function to remove duplicate rows from a DataFrame. If needed, I specify particular columns using the subset parameter to identify duplicates. In ETL pipelines, removing duplicates is an important data quality step before loading data into the target system.
How to explode a nested JSON using Pandas?
click to reveal answerInterview Answer: For nested JSON data, I first flatten the JSON using json_normalize(). If a column contains a list of values, I use the explode() function to convert each list element into a separate row. This makes the data easier to transform and load into relational tables.
What is the difference between Pandas vs PySpark?
click to reveal answerInterview Answer: Pandas processes data on a single machine and is best for small to medium datasets that fit into memory. PySpark is built for distributed computing and can process large-scale data across multiple cluster nodes. In my projects, I use Pandas for local analysis and PySpark for big data ETL pipelines in production.