Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Logging instead of print

Python · Python for Data Pipelines

Logging instead of print

Easypython-45
loggingobservabilityproductionstructured-logs

Question

Why should production pipelines use the logging module rather than print?

Solution

print writes text to the screen and nothing else. The logging module gives each message a severity, a time, a source, and a destination you can change without touching the code. That is what makes logs usable once a job runs on a schedule in a cluster.

What logging adds

import logging

logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s %(levelname)s %(name)s %(message)s",
)
log = logging.getLogger("orders_etl")

log.info("loaded %d rows for %s", row_count, run_date)
log.warning("skipped %d bad rows", bad_count)
log.error("connection failed", exc_info=True)

What you get from it:

  • Levels: DEBUG, INFO, WARNING, ERROR, CRITICAL. You can show only warnings in production and turn on debug output for a single run, without editing code.
  • Timestamps and logger names are added for you, so you know when and where a line came from.
  • Handlers: the same messages can go to the console, a file, or a log service. Production setups send them to Cloud Logging, CloudWatch or Datadog.
  • Tracebacks: log.exception(...) or exc_info=True records the full stack trace.

Structured logs

Plain text lines are hard to search at scale. Many teams emit JSON, one object per line, with fields such as run_id, table, rows and duration_ms. Log tools can then filter on table = "orders" and chart the row counts. Libraries such as structlog or a JSON formatter do this.

A correlation id

Put a run id on every line, either as a field or through a LoggerAdapter. When a job has twenty tasks writing logs in parallel, the run id lets you pull out exactly the lines of one run.

What not to log

Passwords, API keys, tokens, and personal data. Also avoid logging every row or entire payloads: it is slow, expensive to store, and a privacy risk. Log counts, ids and decisions.

Orchestrators

Airflow and similar tools capture whatever your task writes to stdout and the logging output, and show it in the UI. If you use print, you get unlabeled lines with no level, and it is hard to filter. Say in an interview that logging is a basic operational habit, and that you would log what a person on call needs to know at 3 a.m.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext