·7 min readMachine Learning

Vectors and Tensors Explained

A visual guide to how data is shaped in memory: scalars, vectors, matrices, Pandas DataFrames, and multi-dimensional tensors for images, audio, and video.

When you start working with data science, machine learning, or computer vision, you constantly hear terms like scalars, vectors, matrices, and tensors.

Libraries like NumPy, PyTorch, TensorFlow, and Pandas throw shapes like (32, 3, 224, 224) or (1000, 1) at you all the time. But what do these dimensions actually represent in the real world?

How does an audio recording turn into a 1D vector? How does a picture become a 3D tensor? And how does a video become a 4D tensor?

Let’s build the intuition from scratch, dimension by dimension.


The Big Picture: What is a Tensor?

Before diving into code, here is the simplest mental model:

A tensor is just a container for numbers with a specific number of dimensions (rank).

Dimension (Rank)    Math Name     Programming / Data Example
---------------------------------------------------------------------
0D                  Scalar        Single number: 42, 3.14
1D                  Vector        List, Pandas Series, Audio waveform
2D                  Matrix        Pandas DataFrame, Grayscale image
3D                  3D Tensor     RGB Color image (Height, Width, Channels)
4D                  4D Tensor     Video (Frames, Height, Width, Channels)
5D                  5D Tensor     Batch of Videos (Batch, Frames, H, W, C)

Fascinating, isn’t it? Every complex piece of data in machine learning is just numbers organized along different axes.


0D: The Scalar (A Single Point)

A 0D tensor is simply a single number. It has zero dimensions, meaning it has no direction or length.

import numpy as np

temperature = 24.5
scalar = np.array(temperature)

print(scalar.shape)  # Output: () -> rank 0
[ 24.5 ]
  (0D)

Examples of 0D scalars:

  • Your age (e.g. 22)
  • The confidence score of a model (e.g. 0.98)
  • Loss value during training (e.g. 0.042)

1D: Vectors and Pandas Series (A Single Line)

When you line up multiple numbers in a single row or column, you get a 1D tensor, commonly called a vector.

import pandas as pd
import numpy as np

# A NumPy 1D array
vector = np.array([10, 20, 30, 40])
print(vector.shape)  # Output: (4,)

# A Pandas Series is also a 1D vector with labeled indices!
prices = pd.Series([100, 150, 200], index=["Apple", "Banana", "Cherry"])

Visualizing a 1D Vector

Axis 0 (Length: 4)
┌────┬────┬────┬────┐
│ 10 │ 20 │ 30 │ 40 │
└────┴────┴────┴────┘
Shape: (4,)

Real-World 1D Example: Digital Audio

How do you store sound as a 1D vector?

Sound is a continuous wave of air pressure. When a microphone records audio, it measures the amplitude of that wave thousands of times per second (known as the sampling rate):

Waveform:  ~~~/\_/\~~~

Sampled Points:
[ 0.02, 0.15, 0.84, 0.62, -0.21, -0.75, 0.05, ... ]
Shape: (44100,) -> 1 second of mono audio at 44.1 kHz

Each sample is just a number. If you record 5 seconds of mono audio at 44.1 kHz, you get a 1D vector of shape (220500,).


2D: Matrices, Pandas DataFrames, and Grayscale Images

When you arrange numbers in a grid of rows and columns, you get a 2D tensor, traditionally called a matrix.

# A 2D NumPy array
matrix = np.array([
    [1, 2, 3],
    [4, 5, 6]
])
print(matrix.shape)  # Output: (2, 3) -> (Rows, Columns)

Visualizing a 2D Grid

          Axis 1 (Columns)
          Col 0  Col 1  Col 2
Axis 0  ┌──────┬──────┬──────┐
Row 0   │  1   │  2   │  3   │
        ├──────┼──────┼──────┤
Row 1   │  4   │  5   │  6   │
        └──────┴──────┴──────┘
Shape: (2, 3)

1. Pandas DataFrame (Tabular Data)

A Pandas DataFrame is a 2D matrix where columns have names and rows have indices:

df = pd.DataFrame({
    'Age': [25, 30, 35],
    'Salary': [50000, 65000, 80000],
    'Score': [88.5, 92.0, 79.5]
})

print(df.shape)  # Output: (3, 3) -> 3 rows, 3 columns

2. Grayscale Images (Black & White)

How does a black and white image live in memory?

It is literally a 2D matrix of pixel brightness values between 0 (pure black) and 255 (pure white):

Height (Rows) x Width (Columns)

┌─────────────────────────────────┐
│   0    0    0    0    0    0    │  <- Pure Black
│   0  255  255  255  255    0    │  <- White line
│   0  255    0    0  255    0    │  <- White edges
│   0  255  255  255  255    0    │
│   0    0    0    0    0    0    │
└─────────────────────────────────┘
Shape: (5, 6) -> (Height: 5, Width: 6)

In the classic MNIST handwritten digits dataset, every single image is a 2D matrix of shape (28, 28).


3D: Color Images (RGB) and Sequential Data

Now things get interesting. What happens when an image has color?

As we know, colors are made by combining Red, Green, and Blue (RGB). So a color image needs three 2D matrices stacked together, one for each color channel.

# A color image tensor in NumPy / OpenCV
# Shape: (Height, Width, Channels)
image = np.zeros((1080, 1920, 3), dtype=np.uint8)
print(image.shape)  # Output: (1080, 1920, 3)

Visualizing a 3D RGB Tensor

                 ┌───────────────┐
                /   BLUE Layer  /│
               ┌───────────────┐ │
              /  GREEN Layer  /│ │
             ┌───────────────┐ │/│
             │   RED Layer   │ │ │
      Height │   (H x W)     │ │/
             │               │/
             └───────────────┘
                   Width

Shape: (Height, Width, Channels) = (H, W, 3)
  • Channel 0 (Red): Matrix of red intensity values (0-255)
  • Channel 1 (Green): Matrix of green intensity values (0-255)
  • Channel 2 (Blue): Matrix of blue intensity values (0-255)

When you look at pixel (x, y), the computer reads [R, G, B] at that exact coordinate to display the final color.

Note for Deep Learning: In PyTorch, color images are typically shaped as (Channels, Height, Width) or (C, H, W), whereas in OpenCV and TensorFlow they are (Height, Width, Channels) or (H, W, C).


4D: Videos and Batches of Images

How do we take another step up to 4 dimensions?

There are two primary ways a 4D tensor is used in real applications:

1. A Video (Time + 3D Color Image)

What is a video? A video is just a collection of color images (frames) played across time:

Shape: (Frames, Height, Width, Channels)
       └───┬──┘ └───────────┬──────────┘
         Time      3D Color Image

If you have a 10-second video at 30 FPS with 1080p resolution:

  • Frames: 10 seconds * 30 FPS = 300 frames
  • Height: 1080
  • Width: 1920
  • Channels: 3 (RGB)
video_tensor_shape = (300, 1080, 1920, 3)
 Frame 0        Frame 1        Frame 2              Frame N
┌─────────┐    ┌─────────┐    ┌─────────┐         ┌─────────┐
│ RGB (3) │ ── │ RGB (3) │ ── │ RGB (3) │ ──...── │ RGB (3) │
└─────────┘    └─────────┘    └─────────┘         └─────────┘
  (H x W)        (H x W)        (H x W)             (H x W)

  ◄──────────────────── Time / Sequence ────────────────────►

2. A Batch of Images in Machine Learning

When training a neural network (like a ResNet or YOLO), you do not feed one image at a time. You feed a batch of images together to maximize GPU utilization.

Shape: (Batch_Size, Height, Width, Channels)
       (32, 224, 224, 3) -> 32 color images of size 224x224

5D: Batches of Videos in Deep Learning

What if you are training an AI model to classify video actions (like recognizing whether someone is running or jumping)?

You need to feed a batch of multiple videos at once:

Shape: (Batch_Size, Frames, Height, Width, Channels)
       (8, 30, 224, 224, 3)
Dimension Index    Dimension Name     Value in Example
------------------------------------------------------
Axis 0             Batch Size         8 videos
Axis 1             Time (Frames)      30 frames per video
Axis 2             Height             224 pixels
Axis 3             Width              224 pixels
Axis 4             Channels           3 (RGB)

That is 5 dimensions working together in a single structured block of memory.


Summary Cheat Sheet

Here is how all common media and data types map directly to dimensions:

Data Type Tensor Rank Typical Shape Meaning of Axes
Scalar 0D () Single value (e.g. loss, score)
Pandas Series 1D (N,) (Samples,)
Mono Audio 1D (Samples,) (Sample_Rate * Seconds,)
Pandas DataFrame 2D (Rows, Cols) (Records, Features)
Grayscale Image 2D (H, W) (Height, Width)
Stereo Audio 2D (2, Samples) (Channels, Samples)
Color Image (RGB) 3D (H, W, 3) (Height, Width, Channels)
Batch of Grayscale 3D (B, H, W) (Batch, Height, Width)
Single Video 4D (F, H, W, 3) (Frames, Height, Width, Channels)
Batch of Color Img 4D (B, H, W, 3) (Batch, Height, Width, Channels)
Batch of Videos 5D (B, F, H, W, 3) (Batch, Frames, Height, Width, Channels)

Whenever you see a complicated tensor shape in code, just remember: it is always built by stacking simpler dimensions on top of each other.

That’s it folks.