Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Pandas for Data Manipulation

Progress0/20
x

Getting Started with Pandas

  • What is pandas?10m
  • Pandas versus SQL10m
  • Reading CSV, Excel, and JSON10m
  • apply and map12m

Frames & cleaning

  • Series & DataFrames8m
  • Data cleaning & type casting10m
  • Filtering & slicing8m

Transforms & shape

  • Transformations & derived columns10m
  • Grouping & aggregations10m
  • Combining datasets10m
  • Reshaping & pivot tables10m
  • Time series in pandas10m
  • Memory optimization10m

Scaling Up

  • Parquet, chunks, and out-of-core12m
  • Advanced groupby: transform, filter, custom agg12m
  • MultiIndex / hierarchical indexing12m

Production Pandas

  • Data validation & schema contracts12m
  • SQL ⟷ pandas round-trip10m
  • The .str accessor for text cleaning12m
  • Capstone: messy data to analysis-ready16m
Back to track
  1. Learn
  2. Pandas for Data Manipulation
  3. Frames & cleaning
  4. Series & DataFrames

Lesson 5 of 20 · Theory first, then run it

Series & DataFrames

pandasbeginner8 min

Overview

A Series is one labeled column. A DataFrame is those columns standing side by side.

On this page7 sections›
  1. 1The idea
  2. 2Why this exists
  3. 3Picture this
  4. 4A small example
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The idea

Pandas is a Python library for working with tabular data: rows and columns, like a spreadsheet. The two core objects are the Series and the DataFrame. A Series is a single column of values with labels on the left side (called the index). A DataFrame is a collection of Series that share the same index, forming a full table with named columns across the top and labeled rows down the side.

Think of a DataFrame as a spreadsheet tab. Each column has a name (like 'order_id' or 'total'), each row has a label (usually a number starting from 0), and every cell holds one value. A Series is just one of those columns pulled out on its own.

The index is the set of row labels. By default, pandas assigns integer labels starting from 0 (called a RangeIndex). You can replace the default index with meaningful labels like order IDs or dates using set_index. The index controls how pandas aligns data when you combine or compare different objects.

Why this exists

Imagine you work at an online store. Every day, thousands of orders come in as CSV files or database exports. You need to answer questions like 'how many orders were paid?' or 'what is the average order value?' Before you can answer those questions, you need to load the data into a structure that lets you select specific rows and columns, filter by conditions, and compute statistics. A DataFrame gives you that structure.

Without understanding how DataFrames and Series work, you will struggle with every pandas operation that follows. Selecting the wrong rows, confusing labels with positions, or accidentally collapsing a DataFrame into a Series are the most common beginner mistakes. Getting this foundation right saves hours of debugging later.

Picture this

This editor preloads several DataFrames: df_orders, df_events, df_order_items, df_payments, df_telemetry, df_customers, and df_payloads. Assign the frame you want to display to the variable called result.

Before doing anything with a DataFrame, inspect it. Check its shape (how many rows and columns), its column names, and its data types. This tells you the 'grain' of the data: what each row represents. Is one row an order? A line item? An event? That question matters for every operation you will learn.

Anatomy of a DataFrame
indexorder_idstatustotal0O-104paid84.201O-109pending31.002O-116paid112.40

Column labels run across the top. Index labels run down the left side. Each cell holds one value.

A small example

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

PythonInspect df_orders: shape, columns, dtypes, and a preview
print("shape:", df_orders.shape)
print("columns:", df_orders.columns.tolist())
print(df_orders.dtypes)

one_column = df_orders["order_total"]
print(type(one_column))

result = df_orders[["order_id", "order_status", "order_total"]].head(8)

df_orders.shape returns a tuple like (500, 6), meaning 500 rows and 6 columns. df_orders.columns.tolist() gives you the column names as a plain Python list. df_orders.dtypes shows the data type of each column: int64 for whole numbers, float64 for decimals, object for text.

Series vs DataFrame: one column vs many

When you select a single column with df_orders["order_total"], you get a Series (one-dimensional). When you select columns with a list like df_orders[["order_id", "order_total"]], you get a DataFrame (two-dimensional). This distinction matters because some methods work differently on Series vs DataFrames.

Single brackets with one string give a Series. Double brackets (or a list) give a DataFrame.

Selection syntaxReturnsDimensionsExample use
df["col"]Series1DOne column of values
df[["col"]]DataFrame2DOne-column table
df[["a", "b"]]DataFrame2DMulti-column table

The index is a row address, not a row ID

Every DataFrame has an index. The default index looks like row numbers (0, 1, 2, ...), but it is not always a sequence. If you filter rows, the original index labels stay, leaving gaps. If you call set_index("order_id"), the order_id column becomes the index.

Do not assume that index label 5 means the sixth row. After filtering, label 5 might be the third remaining row. Treat business keys (like order_id) as regular columns unless you deliberately promote one to the index and verify it is unique.

PythonPromote order_id to the index and check uniqueness
orders_by_id = df_orders.set_index("order_id", drop=False)
print("unique index:", orders_by_id.index.is_unique)
print("index name:", orders_by_id.index.name)
result = orders_by_id.head(6)

Alignment follows labels

When you do arithmetic between two Series, pandas matches them by index label, not by visible row position. If the indexes do not match, you get NaN values. Reset or validate indexes before combining independently filtered Series.

.loc selects by label

.loc uses labels to select rows and columns. The syntax is df.loc[row_selection, column_selection]. A colon means 'all' on that axis. For example, df.loc[:, ["order_id", "order_total"]] selects all rows and two named columns.

Label slices in .loc are inclusive at both ends. On an index labeled 100 through 110, .loc[100:105] includes label 105. This is different from normal Python slicing, which excludes the endpoint.

PythonSelect paid orders and specific columns with .loc
columns = ["order_id", "user_id", "order_total"]
paid_mask = df_orders["order_status"].eq("paid")
result = df_orders.loc[paid_mask, columns].head(10)

.iloc selects by position

.iloc uses integer positions, like normal Python indexing. The first row is position 0. A slice like :10 stops before position 10, following standard Python rules.

Use .iloc when the requirement says 'the first 10 rows' or 'columns 2 through 4'. Use .loc when the requirement names specific statuses, IDs, or column labels.

Pandas way vs SQL way: .loc is like WHERE + SELECT, .iloc is like LIMIT with OFFSET.

SelectorUnderstandsSlice endpointBest for
.locLabels and boolean masksInclusiveNamed columns, filtered rows
.ilocInteger positionsExclusiveFirst N rows, positional samples
head(N)First N positionsN rowsQuick preview

SQL connection

In SQL, you write SELECT col1, col2 FROM table WHERE status = 'paid'. In pandas, the equivalent is df.loc[df['status'] == 'paid', ['col1', 'col2']]. The .loc operation combines WHERE (row filter) and SELECT (column choice) in one step.

Two selection routes to the same data
df_orders.loc: labels.iloc: positionsresult grid

Labels route through .loc. Physical positions route through .iloc. Both produce a result.

Avoid chained indexing

Writing df[df['status'] == 'paid']['total'] is called chained indexing. It can trigger a SettingWithCopyWarning and sometimes returns a view instead of a copy, leading to confusing bugs. Use one .loc operation instead, and call .copy() before editing a subset.

Chained indexing hides ownership

df[mask][col] may return a view or a copy depending on internal memory layout. Use df.loc[mask, col] in one step, then .copy() if you plan to modify the result.

Copy-paste without reading the output

Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.

Common beginner questions

Why does pandas start counting from 0?

Python uses zero-based indexing. The first element is always at position 0. This is a language convention, not a pandas choice.

What is the difference between df['col'] and df[['col']]?

Single brackets return a Series (one-dimensional). Double brackets return a DataFrame (two-dimensional, even if it has only one column). Use double brackets when downstream code expects a DataFrame.

When should I use set_index?

Use it when you want a meaningful row label (like a date or an ID) and you know that label is unique. For most beginner work, the default integer index is fine.

What comes next

Now that you can load, inspect, and select data from a DataFrame, the next lesson covers cleaning: handling missing values, fixing data types, and normalizing text. Every real dataset has dirt, and you need to clean it before analysis.

Practice

Build a ten-row projection from df_orders. Select order_id, user_id, and order_total by label, then take the first ten physical rows. Assign the resulting DataFrame to result.

Before running, confirm the expected grain: one row per order. The studio grid displays whatever you assign to result.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
apply and mapData cleaning & type casting