🔍 Ctrl+K
🟢 Easy

SCD Types Overview

TypeBehavior
SCD 0Fixed, no updates allowed
SCD 1Overwrite on top
SCD 2New row with start_date, end_date, flag
SCD 3Add new column for previous value
SCD 4Historical data in different table
SCD 6Hybrid (1+2+3)
🔴 Advanced

SCD 2 Example (from file)

idnameaddressstartdateenddateflag
1indumati123 mg road0101202503082025N
1indumati234 mg road0308202531129999Y
Python
from pyspark.sql import SparkSession
from pyspark.sql.functions import lit, current_date
spark = SparkSession.builder.appName('myapplication').getOrCreate()
newdf = customerdf.withColumn('startdate',current_date()).withColumn('enddate',lit('9999-12-31')).withColumn('FLAG',lit('Y'))

SCD 3 example from file: id | name | address | job title | previous job title — tracks previous value in new column.