AWS Interview Q&A
Covers S3, Glue, Lambda, Redshift, DMS, Kinesis, IAM, Data Warehouse vs Data Lake, CI/CD, and more.
Source: GOLDEN_QUESTIONNAIRE_JULY_2026.pdf • Answers are hidden — click a question to reveal its full interview answer. Use bookmarks + Mark as Complete to track prep.
What is Amazon S3? What are the different S3 Storage Classes?
click to reveal answerInterview Answer: Amazon S3 is an object storage service used to store large amounts of structured and unstructured data with high durability and scalability. Common storage classes include Standard, Intelligent-Tiering, Standard-IA, One Zone-IA, Glacier Instant Retrieval, Glacier Flexible Retrieval, and Glacier Deep Archive. I typically store raw and processed data in S3 for ETL pipelines.
What is AWS Glue? What are the components and key features of it?
click to reveal answerInterview Answer: AWS Glue is a serverless ETL service used to discover, transform, and load data. Its main components are the Data Catalog, Crawler, ETL Jobs, Triggers, Workflows, and Job Bookmarks. Key features include automatic schema discovery, serverless execution, Spark-based processing, and easy integration with S3, Redshift, and Athena.
What are AWS Glue Dynamic Frames? How do they differ from Spark DataFrames?
click to reveal answerInterview Answer: Dynamic Frames are AWS Glue's data structure designed for semi-structured and inconsistent data. Unlike Spark DataFrames, they can automatically handle schema variations and missing fields without failing. When I need Spark SQL functions or better performance, I convert Dynamic Frames into DataFrames.
What is the difference between Data Warehouse vs Data Lake vs Data Lakehouse?
click to reveal answerInterview Answer: A Data Warehouse stores structured, curated data for reporting and BI. A Data Lake stores raw structured and unstructured data at low cost. A Data Lakehouse combines the flexibility of a data lake with the reliability and ACID capabilities of a data warehouse, making it suitable for both analytics and machine learning.
What are Glue Job Bookmarks? How do they enable incremental loads?
click to reveal answerInterview Answer: Glue Job Bookmarks track the data that has already been processed in previous job runs. During the next execution, Glue processes only new or modified records instead of the entire dataset. This enables efficient incremental loading, reduces processing time, and lowers costs.
What is the difference between Glue vs EMR vs Lambda? When do you use each?
click to reveal answerInterview Answer: Glue is a serverless ETL service for data integration. EMR is a managed Hadoop and Spark cluster used for large-scale big data processing with more control. Lambda is a serverless compute service for short eventdriven tasks. I use Glue for ETL, EMR for complex Spark workloads, and Lambda for automation and event-based processing.
How do you deploy your Glue code to higher environments (CI/CD)?
click to reveal answerInterview Answer: In my projects, Glue scripts are stored in Git for version control. A CI/CD pipeline using services like CodePipeline, CodeBuild, or Jenkins validates and deploys the code to Dev, QA, and Production environments. Configuration values such as bucket names and IAM roles are managed separately to avoid code changes.
What is AWS Lambda? What are its key features? What are its limitations?
click to reveal answerInterview Answer: AWS Lambda is a serverless compute service that runs code in response to events without managing servers. Key features include automatic scaling, event-driven execution, and pay-per-use pricing. Its limitations include a maximum execution time of 15 minutes, deployment package size limits, memory limits, and concurrency limits.
What are the advantages of Athena?
click to reveal answerInterview Answer: Amazon Athena is a serverless query service that allows SQL queries directly on data stored in S3. It requires no infrastructure management, integrates with the Glue Data Catalog, supports multiple file formats like Parquet and ORC, and charges only for the amount of data scanned. It is ideal for ad hoc analytics.
Explain the architecture of Amazon Redshift.
click to reveal answerInterview Answer: Amazon Redshift follows a Massively Parallel Processing (MPP) architecture. It consists of a Leader Node, which receives SQL queries and creates execution plans, and multiple Compute Nodes, which process data in parallel. Data is stored in a columnar format, enabling high-performance analytical queries.
Difference between Athena vs Redshift vs Redshift Spectrum.
click to reveal answerInterview Answer: Athena queries data directly from S3 without storing it. Redshift is a fully managed data warehouse optimized for high-performance analytics on stored data. Redshift Spectrum allows Redshift to query external data stored in S3 without loading it into Redshift. I use Athena for ad hoc queries, Redshift for BI dashboards, and Spectrum for combining warehouse and data lake data.
How do you optimize Athena query performance?
click to reveal answerInterview Answer: I optimize Athena by storing data in Parquet or ORC format, partitioning large datasets, compressing files, selecting only required columns instead of SELECT *, and avoiding scanning unnecessary partitions. These techniques reduce the amount of data scanned, improve query performance, and lower query costs.
What are the distribution styles in Redshift? What is the syntax for Distribution Key?
click to reveal answerInterview Answer: Redshift supports AUTO, EVEN, KEY, and ALL distribution styles. EVEN distributes rows equally, KEY distributes rows based on a selected column, ALL copies small tables to every node, and AUTO lets Redshift choose the best option. I usually use DISTKEY on frequently joined columns to reduce data movement.
CREATE TABLE employee (
emp_id INT,
dept_id INT,
name VARCHAR(50)
) DISTSTYLE KEY DISTKEY(dept_id);
Dist Key vs Sort Key in Redshift (4 differences)
click to reveal answerInterview Answer: Dist Key determines how data is distributed across compute nodes, while Sort Key determines how data is physically sorted within each node. Dist Key improves join performance by minimizing data movement, whereas Sort Key improves query performance by reducing data scanned. I use Dist Key on join columns and Sort Key on frequently filtered columns like dates.
What is Amazon CloudWatch? What is it used for?
click to reveal answerInterview Answer: Amazon CloudWatch is a monitoring service for AWS resources and applications. It collects metrics, logs, and events, allowing us to monitor system health and performance. I use CloudWatch to monitor Glue jobs, Lambda executions, EC2 instances, and to create alarms for job failures or high resource usage.
What is AWS DMS (Database Migration Service)? How does it capture updates using CDC?
click to reveal answerInterview Answer: AWS DMS is a managed service used to migrate data between databases with minimal downtime. It supports both full data load and Change Data Capture (CDC). CDC continuously reads database transaction logs to capture inserts, updates, and deletes, and replicates only those changes to the target database.
What are SCD Types? Explain SCD Type 2 with an example.
click to reveal answerInterview Answer: Slowly Changing Dimensions (SCD) are techniques used to manage changes in dimension data. Common types are Type 1, Type 2, and Type 3. SCD Type 2 preserves history by creating a new record whenever a value changes, while marking the old record as inactive using effective dates or an active flag.
Before Update
Customer_ID City Start_Date End_Date Active
101 Delhi 2024-01-01 NULL Y
After Update
Customer_ID City Start_Date End_Date Active
101 Delhi 2024-01-01 2025-06-30 N
101 Mumbai 2025-07-01 NULL Y