sys.getsizeof(obj) returns the size of the object itself in bytes (including CPython object overhead for that container), not the size of all objects it references.
Examples of what it misses:
- Elements inside a list/dict
- Attributes on an instance
- Nested structures
- Allocator arena overhead / fragmentation
- Interpreter and module baseline RSS
So a list of a million strings can show a modest getsizeof(list) while process RSS is huge because of the string objects.
import sys
s = "x" * 1000
lst = [s] * 1000 # 1000 references to the SAME string
print(sys.getsizeof(s)) # size of one str object
print(sys.getsizeof(lst)) # size of list object / pointer array
# NOT 1000 * sizeof(s), and sharing means unique payload is one string
nested = {"a": [1, 2, 3], "b": [1, 2, 3]}
print(sys.getsizeof(nested)) # dict shell only
# Rough deep size sketch (not perfect; watch for shared/cyclic refs)
def deep_getsizeof(obj, seen=None):
if seen is None:
seen = set()
obj_id = id(obj)
if obj_id in seen:
return 0
seen.add(obj_id)
size = sys.getsizeof(obj)
if isinstance(obj, dict):
size += sum(deep_getsizeof(k, seen) + deep_getsizeof(v, seen) for k, v in obj.items())
elif isinstance(obj, (list, tuple, set, frozenset)):
size += sum(deep_getsizeof(i, seen) for i in obj)
return sizeFor process-level usage, inspect RSS via OS tools (ps, resource, tracemalloc, memory profilers), not getsizeof alone.