Azure Interview Q&A
Covers Azure Data Factory, Databricks, Delta Lake, Unity Catalog, managed vs external tables, PII security, and more.
Source: GOLDEN_QUESTIONNAIRE_JULY_2026.pdf • Answers are hidden — click a question to reveal its full interview answer. Use bookmarks + Mark as Complete to track prep.
Why did you use Azure Data Factory (ADF) in the project?
click to reveal answerInterview Answer: We used Azure Data Factory to automate and orchestrate our ETL pipeline. It helped us move data from different source systems to Azure Data Lake and Databricks. ADF also handled scheduling, monitoring, and error handling, reducing manual effort.
How did you implement incremental loading in ADF?
click to reveal answerInterview Answer: We implemented incremental loading using a Watermark column such as Last_Updated_Date. During each pipeline run, ADF loaded only the new or updated records instead of the entire dataset. This reduced processing time and improved performance.
Which ADF activities were used?
click to reveal answerInterview Answer: In my project, I used Copy Data Activity to move data, Lookup Activity to read configuration values, ForEach Activity to process multiple files, If Condition for validations, Stored Procedure Activity for database operations, and Execute Pipeline to call another pipeline.
What is Unity Catalog?
click to reveal answerInterview Answer: Unity Catalog is Databricks' centralized data governance solution. It helps manage tables, files, permissions, and data access from one place. It provides fine-grained security, auditing, and easy data sharing across teams.
What is Liquid Clustering?
click to reveal answerInterview Answer: Liquid Clustering is a Databricks feature that automatically organizes data without manually repartitioning it. It improves query performance by reducing the amount of data scanned. It is more flexible than traditional partitioning because data can be reorganized automatically.
What is Delta Lake? What are its features?
click to reveal answerInterview Answer: Delta Lake is an open-source storage layer built on top of a data lake. It provides ACID transactions, schema enforcement, schema evolution, time travel, and data versioning. These features make data more reliable and suitable for ETL pipelines.
What is the difference between Managed Tables and External Tables?
click to reveal answerInterview Answer: Managed Tables store both data and metadata inside Databricks. If the table is deleted, the data is also deleted. External Tables store only the metadata in Databricks, while the actual data remains in external storage like ADLS or S3. Deleting the table does not delete the data.
How does Azure Key Vault help?
click to reveal answerInterview Answer: Azure Key Vault securely stores secrets such as passwords, API keys, and connection strings. Instead of hardcoding credentials in the code, applications retrieve them securely from Key Vault. This improves security and simplifies secret management.
What is Databricks and what is the Architecture of Databricks?
click to reveal answerInterview Answer: Azure Databricks is a cloud-based platform for big data processing and machine learning using Apache Spark. Its architecture includes a Control Plane, which manages clusters and notebooks, and a Data Plane, where Spark clusters process the data. This architecture provides scalability, performance, and easy collaboration.
How do you secure PII data?
click to reveal answerInterview Answer: I secure PII data by masking or encrypting sensitive columns such as names, email IDs, and phone numbers. I use role-based access control so only authorized users can view the data. I also store secrets in Azure Key Vault and enable encryption for data at rest and in transit.