Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Pandas for Data Manipulation

Progress0/20
x

Getting Started with Pandas

  • What is pandas?10m
  • Pandas versus SQL10m
  • Reading CSV, Excel, and JSON10m
  • apply and map12m

Frames & cleaning

  • Series & DataFrames8m
  • Data cleaning & type casting10m
  • Filtering & slicing8m

Transforms & shape

  • Transformations & derived columns10m
  • Grouping & aggregations10m
  • Combining datasets10m
  • Reshaping & pivot tables10m
  • Time series in pandas10m
  • Memory optimization10m

Scaling Up

  • Parquet, chunks, and out-of-core12m
  • Advanced groupby: transform, filter, custom agg12m
  • MultiIndex / hierarchical indexing12m

Production Pandas

  • Data validation & schema contracts12m
  • SQL ⟷ pandas round-trip10m
  • The .str accessor for text cleaning12m
  • Capstone: messy data to analysis-ready16m
Back to track
  1. Learn
  2. Pandas for Data Manipulation
  3. Getting Started with Pandas
  4. What is pandas?

Lesson 1 of 20 · Theory first, then run it

What is pandas?

pandasbeginner10 min

Overview

Pandas is a Python library for tables. A DataFrame is the table. A Series is one column.

On this page7 sections›
  1. 1The idea
  2. 2Why this exists
  3. 3Picture this
  4. 4A small example
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The idea

Pandas is a Python library for working with tables. If you have opened a spreadsheet, you already know the shape: rows going down, named columns going across, and one value in each cell.

The two objects you will use every day are the DataFrame and the Series. A DataFrame is the full table. A Series is one column of that table, with a label on every row.

At work you will see import pandas as pd at the top of almost every notebook. In this tab you do not need to import pandas or open a file. The studio already built DataFrames named df_orders, df_events, df_order_items, df_payments, df_telemetry, df_customers, df_payloads, and df_orders_raw. Assign the table you want to show to a variable called result.

Before you start

This track assumes you completed Core Python. If you have not, start there first so lists, loops, and functions feel familiar.

Why this exists

An online store collects orders, payments, and click events every day. Those files arrive as CSV exports, Excel workbooks, or JSON dumps. A spreadsheet can open a small file so you can glance at it. Spreadsheets struggle when you need the same cleanup every morning, a join between two tables, or a result that another Python script can use.

Data engineers put that work in pandas so the steps are code, not a sequence of clicks. You inspect a file, pick columns, filter rows, and hand a clean table to the next step. Without a table object in Python you would loop over lists of dictionaries by hand. That is slow to write and easy to get wrong.

Picture this

Three places a table can live
Spreadsheet tabpandas DataFrame (this editor)SQL table in a warehouse

Same rows and columns. Different homes: a spreadsheet file, a pandas DataFrame in Python, or a table in a warehouse.

Think of a DataFrame as a spreadsheet tab that lives inside Python. Each column has a name such as order_id. Each row has a label called the index, usually 0, 1, 2. Each cell holds one value.

Pandas sits between a spreadsheet and a warehouse: code you can rerun, on data that still fits in memory.

ToolWhat it isBest forWeak at
SpreadsheetA file you click throughA quick look at a small extractRepeatable joins and daily cleanup
pandas DataFrameA table in Python memoryFiles, notebooks, Python-side cleaningData bigger than laptop RAM
SQL tableA table in a warehouseShared, large, governed dataAd-hoc file wrangling on a laptop

Before you change anything, look at the table. Shape tells you how many rows and columns. Column names tell you what you can select. A few sample rows tell you the grain: one row per order, or one row per line item.

A DataFrame is a table
indexorder_idorder_statusorder_total0O-104paid84.201O-109pending31.002O-116paid112.40

Column names across the top. Index labels down the left. Each cell holds one value.

A small example

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

PythonInspect df_orders, then preview the first rows
print("shape:", df_orders.shape)
print("columns:", df_orders.columns.tolist())
print("one column type:", type(df_orders["order_status"]))
result = df_orders.head(6)

df_orders.shape is a pair: (rows, columns). df_orders.columns.tolist() returns the names as a Python list. df_orders["order_status"] is a Series. Passing a list of names, as in df_orders[["order_id", "order_status"]], keeps a DataFrame.

PythonSelect two columns by label and keep eight rows
result = df_orders.loc[:, ["order_id", "order_status"]].head(8)
print(result.columns.tolist())
print(len(result))

.loc talks in labels. The colon on the left means all rows. The list on the right names the columns. .head(8) keeps the first eight rows so the grid stays small. The studio displays whatever DataFrame you assign to result.

result must be a DataFrame

The grid only shows a table. If you assign a Series, a number, or a printed string to result, checks that expect columns will fail. Select with a list of column names (or .loc[:, [...]]) so you keep a DataFrame.

SQL connection

SELECT order_id, order_status FROM orders LIMIT 8 is the warehouse version of df_orders.loc[:, ['order_id', 'order_status']].head(8). Same idea: pick columns, then cap the row count.

Copy-paste without reading the output

Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.

Common beginner questions

Why can't I just use Excel?

You can, for a one-off look at a small file. Excel does not version well, does not join large tables cleanly, and does not plug into Python pipelines. Pandas is the same table idea, written as code you can rerun.

Do I need to install pandas in this tab?

No. The editor already has pandas and the studio frames. On your laptop you would run pip install pandas and then import pandas as pd.

What is the difference between a Series and a DataFrame?

A Series is one column with an index. A DataFrame is several columns sharing that index. Methods that expect a table (several named columns) need a DataFrame. Methods that expect a single list of values work on a Series.

What comes next

The next lesson compares pandas to SQL. You will see the same operations (filter, group, join) in both languages, and when to use each.

Practice

Run Sample to inspect df_orders. Then complete Exercise: keep order_id and order_status, take the first 8 rows, and store that DataFrame in result.

Confirm the grain before you run: one row per order. The grid shows whatever you assign to result.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Pandas for Data ManipulationPandas versus SQL