Python mainly frees memory by reference counting: each object counts how many names or containers refer to it, and when the count reaches zero the memory is released at once. A second mechanism, the cyclic garbage collector, finds groups of objects that refer to each other but are no longer reachable from anywhere else.
Reference counting
import sys a = [] sys.getrefcount(a) # the count includes the temporary reference from the call b = a # now two names refer to the same list del a # removes the name 'a', not the object b = None # last reference gone: the list is freed immediately
del x only removes the name x. The object is destroyed only when nothing refers to it any more. That is why del big_df does not always free a DataFrame: if another variable, a cache, a closure or a notebook output history (_, Out[12]) still holds it, the memory stays.
Reference cycles
If object A holds a reference to B and B to A, their counts never reach zero even when your code has dropped all other references. Common examples: a parent and child that point to each other, an object that stores a bound method of itself in a list, or an exception traceback that references the frame that holds the exception.
class Node:
def __init__(self): self.other = None
a, b = Node(), Node()
a.other, b.other = b, a # cycle
del a, b # counts stay above zero; the cyclic GC must collect themThe cyclic garbage collector runs from time to time, looks for such unreachable groups, and frees them. It organises objects into three generations, scanning young ones often and older ones rarely, because most objects die young.
Where this matters in data jobs
- Long-running workers (streaming consumers, Airflow tasks that loop for hours) can grow in memory over time because of leaks: a global list or dictionary that only ever grows (a cache with no limit), callbacks that are registered and never removed, or large objects kept alive by a cycle.
- Notebooks: old results stay referenced through the output history, so memory builds up until you restart.
What to do
- Find the growth with
tracemalloc,objgraphor a memory profiler, instead of guessing. - Bound caches (
functools.lru_cache(maxsize=...)), and clear collections you no longer need. - Use
weakreffor back-references that should not keep the target alive. - Process large data in chunks and let each chunk go out of scope.
- Call
gc.collect()only for diagnosis or at a deliberate checkpoint, and rarely otherwise. Freed memory also is not always returned to the operating system, so the process size may not shrink right away.
How to answer
Name reference counting first, then the cycle collector and generations, then give the "del removes a name" point and a long-running job example.