Skip to main content

22 docs tagged with "pandas"

View all tags

Aggregation Functions

Aggregation combines multiple values into a single summary value. Common operations include sum, mean, count, min, max, and custom aggregations.

Apply and Map

pandas provides several methods to apply custom functions to data:

Binning and Categorical Data

Binning converts continuous data into discrete intervals (bins). Categorical data represents discrete categories with limited unique values. Both are essential for analysis and visualization.

Boolean Filtering

Boolean filtering selects rows based on conditions. It's one of the most common operations in data

Data Types

Understanding and managing data types is crucial for:

DateTime Basics

Working with dates and times is essential for time series analysis. pandas provides powerful datetime functionality through:

Handling Duplicates

Duplicate rows are common in real-world data. _pandas_ provides simple methods to find and remove them.

Interview Snippets

Common pandas questions asked in data science and analytics interviews. Each snippet includes the problem, solution, and explanation.

loc vs iloc

_pandas_ provides two main indexers for selecting data:

Missing Data

Missing data appears as NaN (Not a Number), None, or NaT (Not a Time) in pandas. Handling it correctly is crucial for data analysis.

Outliers and Data Validation

Outliers are data points that differ significantly from other observations. Data validation ensures your data meets expected criteria before analysis.

Reshaping Data

Reshaping transforms data between wide and long formats. Common operations:

Selecting Columns

Column selection is one of the most common operations in _pandas_. There are several ways to

Sorting and Ranking

Sorting organizes data by values. Ranking assigns positions based on values. Both are essential for analysis and presentation.

String Operations

pandas provides string methods through the .str accessor. These methods work on Series containing strings and are essential for text data cleaning.