Overview
git status names three areas. git add copies a file into the index. git commit snapshots the index. This tab has no Git remote.
On this page7 sections
Goal
Git is a version control system. It records snapshots of your project so you can go back to any earlier version, see what changed, and collaborate without overwriting each other's work. Every data engineering team uses Git to track pipeline code, SQL queries, configuration files, and documentation.
Git organizes your project into three areas. The working tree is the files on disk (ingest.py as you saved it). The index (also called the staging area) is the next snapshot you are assembling. A commit is a recorded snapshot of the index, tagged with a message. Understanding these three areas is the key to reading git status output.
A commit is not the same as saving a file. Saving updates the working tree. A commit permanently records a version in the project history. You choose which changes go into a commit by staging them first with git add, then recording them with git commit.
Why this order
Imagine your team deploys a data pipeline every week. On Thursday, the nightly ingest breaks. Without Git, nobody can answer: what changed since last week? Who changed it? Was the config file different before? Git answers all of these questions because every commit is a timestamped snapshot with an author and a message.
git status reports which files sit in which area: untracked, modified in the working tree, staged in the index, or already committed. Recording a file you just wrote is two commands in order: git add, then git commit. Status after a clean record should show nothing left to stage for that file.
Beyond history, Git enables code review. A teammate can read your commit before it reaches production. Without version control, changes happen silently and mistakes compound for days before anyone notices.
The checklist
The exercises in this tab use Python lists of command strings, not a live Git repository. There is no .git directory and no remote. Real git lives in a terminal on your machine, on a repo you own. After passing, install Git and practice on a real folder.
Mental model: working tree, index, commit
The working tree is files on disk: ingest.py as you saved it, including edits you have not told Git about. The index (staging area) is the next snapshot you are assembling. git add copies a path from the working tree into the index. git commit writes a snapshot of the index, with a message, into the history. HEAD points to the current commit.
git status prints the map. Untracked or modified lives in the working tree. Staged lives in the index. Nothing to commit means the index and HEAD agree for the paths Git is watching. Recording a new ingest.py is therefore add, then commit. Status after that should be quiet for that file.
add copies tree into index. commit snapshots the index. Status reports leftovers.
What each command is for
Status first. Add. Commit. Push is homework on a repo you own.
| Command | Moves what | Leaves what |
|---|---|---|
| git status | Nothing. Reads the three areas. | A report: untracked, staged, committed. |
| git add ingest.py | Working tree bytes into the index | Disk still has the file; history does not, yet. |
| git commit -m "ingest v1" | Index into a commit | A snapshot teammates can fetch later. |
| git push (not here) | Commits to a remote | Needs a remote. This tab has none. |
Worked example: two commands, in order
ingest.py is on disk in the working tree. You do not commit the whole laptop. You add the path, then commit with a message future readers can use. The exercise stores that pair as a Python list.
result = ["git add ingest.py", 'git commit -m "ingest v1"']
print(result)The list order matters. git commit before git add would commit whatever the index already held, not the new file. Always add, then commit.
Checking status before and after
A good habit is running git status before you add and after you commit. The function below simulates checking whether a file appears in the staged list.
def is_staged(file_name, staged_files):
return file_name in staged_files
staged = ["ingest.py"]
print(is_staged("ingest.py", staged))
print(is_staged("utils.py", staged))Lists and loops from Core Python
You practiced list operations in Core Python when you stored and iterated over items. The same skills apply here: commands are strings in a list, and order determines the outcome.
git commit -a skips the staging step
The -a flag adds all tracked, modified files and commits in one step. Beginners use this to skip staging, but it can include files you did not intend to commit. Prefer explicit git add for each path, then commit.
Wrong and right
The wrong habit is git commit -a on files you have not read, or committing after a save and assuming the remote updated. The right habit is status, add the paths you intend, commit a sentence, status again. Push when a remote exists.
- Working tree is disk. Index is the next snapshot. Commit is history.
- git status is the map between those three.
- Add then commit records a file. Push is later, on your machine.
Before you start
This track assumes you completed Core Python. Exercises use lists, dicts, and sorting patterns you already practiced there.
Worked pass
ingest.py sits in the working tree. Recording it is two commands in order: git add, then git commit. The list order is the whole lesson.
Input: a new file you just saved.
| area | ingest.py state |
|---|---|
| working tree | file exists on disk, not yet in history |
| index (after add) | staged for the next commit |
| commit (after commit) | recorded in history |
result = ["git add ingest.py", 'git commit -m "ingest v1"']
print(result)stdout shows add before commit. Reversing the list would commit whatever was already staged, not the new file.
Output: the two-command sequence the exercise grades.
| step | command | effect |
|---|---|---|
| 1 | git add ingest.py | copy path into the index |
| 2 | git commit -m "ingest v1" | snapshot the index into history |
Copy-paste without reading the output
Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.
Common beginner questions
Do I need Git for data engineering?
Yes. Every team stores pipeline code, SQL, and config in a Git repository. Knowing status, add, and commit is the minimum. You will also need branching and pull requests on the job.
What is the difference between git add and git commit?
git add moves your changes into the staging area. git commit records a snapshot of everything staged. Think of add as packing a box and commit as sealing and labeling it.
Can I undo a commit?
Yes. git revert creates a new commit that undoes an earlier one. git reset moves the HEAD pointer. Both require a real .git directory, which this tab does not have.
What comes next
The next lesson covers merge conflicts. When two people edit the same lines, Git stops and asks you to resolve the overlap. Knowing the three areas from this lesson makes conflicts easier to understand.
Practice
Run Sample and read the two-command list. Then complete Exercise: put git add ingest.py and git commit -m "ingest v1" in that order, store the list in result.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.