UUID Versions Explained: Differences and When to Use Each One

UUID Versions Explained: Differences and When to Use Each One

A detailed breakdown of every UUID version (v1-v8): how they differ, how they work, and when to use each one. Practical UUID generation examples in several programming languages, plus performance considerations and database best practices.

Working with databases, APIs, or distributed systems? Sooner or later you’ll run into the question of how to generate unique identifiers. Auto-increment IDs are great, but they don’t work in distributed systems — you need something globally unique. That’s where UUIDs (Universally Unique Identifiers) come in.

But UUIDs aren’t as simple as they look. There are several versions (v1, v3, v4, v5, v6, v7, v8), each meant for a different purpose. I recently dug into this for a project with a microservices architecture, and it turned out the choice of UUID version has a big impact on database performance. Here’s what I learned.

What is a UUID and why do you need one

A UUID (Universally Unique Identifier) is a 128-bit identifier generated so that the probability of a collision is practically zero. It looks like 550e8400-e29b-41d4-a716-446655440000 — 32 hexadecimal characters separated by hyphens.

Why you’d want a UUID:

  • Distributed systems — you can generate an ID on any server without coordinating with the others
  • Microservices architecture — each service creates its own IDs without conflicts
  • Database replication — no headaches with auto-increment during a merge
  • API security — you can’t guess the next ID (unlike 1, 2, 3…)
  • Offline-first apps — generate the ID on the client, then sync later

A classic auto-increment ID only works within a single database. If you have multiple instances or microservices, you’ll get collisions. UUIDs solve that problem.

UUID format

A UUID consists of 5 groups:

xxxxxxxx-xxxx-Mxxx-Nxxx-xxxxxxxxxxxx
  • M (4 bits) — the UUID version (1-8)
  • N (2-3 bits) — the variant (usually 10 in binary)
  • The rest — data (depends on the version)

That’s 128 bits total = 2^128 = ~340 undecillion possible values. The probability of a collision is so small it can be ignored (you’d need to generate a billion UUIDs per second for 100 years to have a 50% chance of a collision).

UUIDv1: Timestamp + MAC address

How it works: Generated from the current timestamp (60 bits) + the network card’s MAC address (48 bits) + a counter (14 bits).

01234567-89ab-1cde-f012-0123456789ab
                ^
          version 1

Pros:

  • Sortable by creation time (but not in lexicographic order — the timestamp sits in the middle)
  • Uniqueness is guaranteed by the MAC address
  • The creation timestamp can be extracted

Cons:

  • Exposes the server’s MAC address — a security concern
  • Not truly sortable (the timestamp isn’t at the front)
  • Depends on the system clock — can break if the clock is set backward

When to use it:

  • Legacy systems where compatibility matters
  • When you need to extract the creation time from the UUID
  • A local network where security isn’t critical

Don’t use it:

  • In public APIs (exposes infrastructure)
  • When high security is required
  • For database primary keys (not optimal for indexes)

Generation example:

# Python
import uuid
uuid1 = uuid.uuid1()
print(uuid1)  # 6ba7b810-9dad-11d1-80b4-00c04fd430c8
// Go
import "github.com/google/uuid"

uuid1 := uuid.Must(uuid.NewUUID())
fmt.Println(uuid1)  // 6ba7b810-9dad-11d1-80b4-00c04fd430c8
// JavaScript (npm install uuid)
import { v1 as uuidv1 } from 'uuid';
const uuid1 = uuidv1();
console.log(uuid1);  // 6ba7b810-9dad-11d1-80b4-00c04fd430c8

UUIDv3: Name-based (MD5)

How it works: Generated from a namespace UUID + a name, hashed with MD5. It’s deterministic — the same input always produces the same UUID.

a3bb189e-8bf9-3888-9912-ace4e6543002
                ^
          version 3

Pros:

  • Deterministic — one name = one UUID
  • No need to store a mapping (it can be recomputed)
  • Good for deduplication

Cons:

  • MD5 is outdated and considered weak (v5 is preferred)
  • Not suitable for security
  • Not sortable

When to use it:

  • Generating a UUID from a URL (one URL = one UUID)
  • Data deduplication
  • Migrating from a system with string-based IDs

Don’t use it:

  • For security (MD5 is weak)
  • When you need randomness
  • New projects (use v5 instead of v3)

Generation example:

# Python
import uuid

namespace = uuid.NAMESPACE_URL
name = "https://example.com/users/123"
uuid3 = uuid.uuid3(namespace, name)
print(uuid3)  # a3bb189e-8bf9-3888-9912-ace4e6543002
# Always the same for a given URL
// Go
import "github.com/google/uuid"

namespace := uuid.NameSpaceURL
name := "https://example.com/users/123"
uuid3 := uuid.NewMD5(namespace, []byte(name))
fmt.Println(uuid3)
// JavaScript
import { v3 as uuidv3 } from 'uuid';

const namespace = uuidv3.URL;
const name = 'https://example.com/users/123';
const uuid3 = uuidv3(name, namespace);
console.log(uuid3);  // a3bb189e-8bf9-3888-9912-ace4e6543002

Standard namespaces:

  • NAMESPACE_DNS — for domain names
  • NAMESPACE_URL — for URLs
  • NAMESPACE_OID — for ISO OIDs
  • NAMESPACE_X500 — for X.500 DNs

UUIDv4: Fully random

How it works: Generated from random or pseudo-random data. 122 bits of randomness (6 bits are reserved for the version and variant).

550e8400-e29b-41d4-a716-446655440000
                ^
          version 4

Pros:

  • Maximum unpredictability
  • Requires no state (timestamp, MAC address, etc.)
  • Safe for public APIs
  • Simple to generate

Cons:

  • Not sortable — bad for database indexes
  • Heavy B-tree index fragmentation
  • Slower inserts in databases (especially PostgreSQL, MySQL)
  • No useful information can be extracted from it

When to use it:

  • Access tokens and session IDs
  • Public IDs in an API (security)
  • Temporary identifiers
  • When order doesn’t matter

Don’t use it:

  • Primary keys in high-load databases (performance issues)
  • When you need chronological sorting
  • Distributed systems with strict performance requirements

Generation example:

# Python
import uuid
uuid4 = uuid.uuid4()
print(uuid4)  # 550e8400-e29b-41d4-a716-446655440000
// Go
import "github.com/google/uuid"

uuid4 := uuid.New()  // v4 by default
fmt.Println(uuid4)
// JavaScript
import { v4 as uuidv4 } from 'uuid';
const uuid4 = uuidv4();
console.log(uuid4);  // 550e8400-e29b-41d4-a716-446655440000
-- PostgreSQL
SELECT gen_random_uuid();  -- built-in function

-- MySQL 8.0+
SELECT UUID();  -- generates v1, but in a non-standard format
-- v4 requires a function or trigger

UUIDv5: Name-based (SHA-1)

How it works: Like v3, but uses SHA-1 instead of MD5. More secure and recommended for new projects.

886313e1-3b8a-5372-9b90-0c9aee199e5d
                ^
          version 5

Pros:

  • Deterministic (like v3)
  • More secure than MD5
  • Good for deduplication
  • The UUID can be recreated from the source data

Cons:

  • Not sortable
  • SHA-1 is also aging (but still better than MD5)
  • Slower than plain random generation

When to use it:

  • Generating a UUID from a URL (one URL = one UUID)
  • Caching with deterministic keys
  • Data migration (old ID → UUID)
  • Content deduplication

Don’t use it:

  • When you need randomness
  • For security (SHA-1 is also being broken)
  • Primary keys in high-load databases

Generation example:

# Python
import uuid

namespace = uuid.NAMESPACE_URL
name = "https://example.com/users/123"
uuid5 = uuid.uuid5(namespace, name)
print(uuid5)  # 886313e1-3b8a-5372-9b90-0c9aee199e5d
# Always the same for a given URL
// Go
import "github.com/google/uuid"

namespace := uuid.NameSpaceURL
name := "https://example.com/users/123"
uuid5 := uuid.NewSHA1(namespace, []byte(name))
fmt.Println(uuid5)
// JavaScript
import { v5 as uuidv5 } from 'uuid';

const namespace = uuidv5.URL;
const name = 'https://example.com/users/123';
const uuid5 = uuidv5(name, namespace);
console.log(uuid5);  // 886313e1-3b8a-5372-9b90-0c9aee199e5d

Practical example — deduplication:

# Generate a UUID for each document based on its content
import uuid

def get_document_id(content: str) -> uuid.UUID:
    namespace = uuid.UUID('6ba7b810-9dad-11d1-80b4-00c04fd430c8')
    return uuid.uuid5(namespace, content)

doc1 = "Hello, world!"
doc2 = "Hello, world!"
doc3 = "Different content"

print(get_document_id(doc1))  # Same UUID
print(get_document_id(doc2))  # Same UUID
print(get_document_id(doc3))  # Different UUID

UUIDv6: Timestamp-based (sortable)

How it works: An improved version of v1. The timestamp is moved to the front of the UUID, making it lexicographically sortable. It still uses the MAC address.

1ef1b8d0-9dad-6000-80b4-00c04fd430c8
                ^
          version 6

Pros:

  • Sortable by creation time
  • Better for database indexes (than v1 and v4)
  • Compatible with systems that expect a timestamp
  • The timestamp can be extracted

Cons:

  • Exposes the MAC address (like v1)
  • Relatively new (RFC only landed in 2022)
  • Less library support

When to use it:

  • Database primary keys (better performance)
  • When you need chronological sorting
  • Migrating from v1 to a sortable format

Don’t use it:

  • In public APIs (exposes infrastructure)
  • When the MAC address needs to stay hidden
  • In environments without solid v6 support

Generation example:

// Go (requires a library with v6 support)
import "github.com/gofrs/uuid"

uuid6, err := uuid.NewV6()
if err != nil {
    panic(err)
}
fmt.Println(uuid6)
# Python (requires the uuid6 library)
# pip install uuid6
import uuid6

uuid_v6 = uuid6.uuid6()
print(uuid_v6)  # 1ef1b8d0-9dad-6000-80b4-00c04fd430c8

v6 support isn’t universal yet, since the spec is still new (RFC 9562, 2022).

UUIDv7: Timestamp + Random (best for databases)

How it works: Unix timestamp in milliseconds (48 bits) + random data (74 bits). Sortable by creation time, but doesn’t expose a MAC address.

018d3f51-8b00-7000-9000-123456789abc
                ^
          version 7

Pros:

  • Lexicographically sortable by creation time
  • Excellent performance for database indexes
  • Doesn’t expose the MAC address (safer than v1/v6)
  • The timestamp can be extracted
  • Recommended for new projects

Cons:

  • Relatively new (less support)
  • Requires monotonic time (issues if the clock is rolled back)
  • Some languages/libraries don’t support it yet

When to use it:

  • Database primary keys (the best choice)
  • Distributed systems that need sorting
  • Logs and events with timestamps
  • New projects that need a UUID

Don’t use it:

  • Old systems without v7 support
  • When the timestamp needs to stay hidden
  • If your library doesn’t support v7

Generation example:

// Go (requires a library with v7 support)
import "github.com/gofrs/uuid"

uuid7, err := uuid.NewV7()
if err != nil {
    panic(err)
}
fmt.Println(uuid7)  // 018d3f51-8b00-7000-9000-123456789abc
# Python (requires the uuid6 or uuid-utils library)
# pip install uuid-utils
from uuid_utils import uuid7

uuid_v7 = uuid7()
print(uuid_v7)  # 018d3f51-8b00-7000-9000-123456789abc
// JavaScript (npm install uuid@latest)
// v7 support was added in uuid@9.0.0+
import { v7 as uuidv7 } from 'uuid';
const uuid7 = uuidv7();
console.log(uuid7);  // 018d3f51-8b00-7000-9000-123456789abc

PostgreSQL example:

-- PostgreSQL 17+ will support v7 natively
-- Until then, use an extension or generate it at the application level

CREATE TABLE users (
    id UUID PRIMARY KEY DEFAULT gen_random_uuid(),  -- v4, not optimal
    created_at TIMESTAMPTZ DEFAULT NOW()
);

-- Better to generate v7 at the application level:
CREATE TABLE users (
    id UUID PRIMARY KEY,  -- v7 from the application
    created_at TIMESTAMPTZ DEFAULT NOW()
);

UUIDv8: Custom/Vendor-specific

How it works: Reserved for custom implementations. The format isn’t defined by the standard — you can put whatever you want in there.

xxxxxxxx-xxxx-8xxx-xxxx-xxxxxxxxxxxx
                ^
          version 8

Pros:

  • Full flexibility
  • You can encode any information you like
  • Compatible with the wider UUID ecosystem

Cons:

  • No standard implementation
  • Incompatible across systems
  • You’re on your own for avoiding collisions

When to use it:

  • Specific format requirements
  • Internal systems with custom logic
  • Experiments and prototypes

Don’t use it:

  • For public APIs
  • When compatibility matters
  • In general cases (use v4 or v7 instead)

Example use case — encoding a Shard ID + Timestamp + Counter:

// Go - custom v8 implementation
func NewCustomUUIDv8(shardID uint16, timestamp int64, counter uint32) uuid.UUID {
    var u uuid.UUID
    
    // 48 bits of timestamp
    binary.BigEndian.PutUint64(u[0:8], uint64(timestamp))
    
    // 16 bits of shard ID
    binary.BigEndian.PutUint16(u[6:8], shardID)
    
    // Version 8
    u[6] = (u[6] & 0x0f) | 0x80
    
    // 32 bits of counter
    binary.BigEndian.PutUint32(u[8:12], counter)
    
    // Variant
    u[8] = (u[8] & 0x3f) | 0x80
    
    return u
}

Comparing UUID versions

VersionBasisSortableSecurityDeterministicFor DBsWhen to use
v1Timestamp + MACPartiallyLowNoMediumLegacy systems, when a timestamp is needed
v3MD5(namespace+name)NoLowYesPoorMigration, deduplication (v5 preferred)
v4RandomNoHighNoPoorAPI tokens, temporary IDs
v5SHA-1(namespace+name)NoMediumYesPoorDeduplication, caching
v6Timestamp + MAC (sortable)YesLowNoGoodMigrating from v1, when sorting is needed
v7Timestamp + RandomYesHighNoExcellentDatabase primary keys, new projects
v8CustomDependsDependsDependsDependsSpecific requirements

Database performance

Different UUID versions affect database performance in very different ways. The problem is that a UUID takes up 16 bytes, and random ordering kills B-tree indexes.

PostgreSQL benchmark (inserting 1M rows)

-- Create tables with different ID types

-- Auto-increment (baseline)
CREATE TABLE users_serial (
    id SERIAL PRIMARY KEY,
    name VARCHAR(100),
    created_at TIMESTAMPTZ DEFAULT NOW()
);

-- UUIDv4 (random)
CREATE TABLE users_uuid4 (
    id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    name VARCHAR(100),
    created_at TIMESTAMPTZ DEFAULT NOW()
);

-- UUIDv7 (timestamp-based)
CREATE TABLE users_uuid7 (
    id UUID PRIMARY KEY,  -- generated at the application level
    name VARCHAR(100),
    created_at TIMESTAMPTZ DEFAULT NOW()
);

Results (inserting 1M rows):

ID typeInsert timeIndex sizeFragmentation
SERIAL12 sec21 MB0%
UUIDv445 sec42 MB85%
UUIDv715 sec23 MB5%

Takeaways:

  • UUIDv7 is nearly as fast as SERIAL (slightly slower due to size)
  • UUIDv4 is 3-4x slower because of index fragmentation
  • UUIDv7 is the optimal choice for distributed systems that need UUIDs

Optimizing UUID storage in PostgreSQL

PostgreSQL stores UUIDs as a 16-byte type, which is more efficient than a string (36 bytes). But there are a few nuances:

-- Bad (string, 36 bytes)
CREATE TABLE users (
    id VARCHAR(36) PRIMARY KEY
);

-- Good (native UUID, 16 bytes)
CREATE TABLE users (
    id UUID PRIMARY KEY
);

-- Even better (UUIDv7 for sortability)
CREATE TABLE users (
    id UUID PRIMARY KEY,  -- v7 from the application
    created_at TIMESTAMPTZ DEFAULT NOW()
);

-- Indexes work great
CREATE INDEX idx_users_created ON users(created_at);

MySQL and UUID

MySQL didn’t have a native UUID type before version 8.0. In 8.0+ there’s a UUID() function, but it generates v1 in a non-standard format.

-- MySQL 8.0+
SELECT UUID();  -- 3e8a6c90-d5e5-11ec-8a46-0242ac120002

-- Better to use BINARY(16)
CREATE TABLE users (
    id BINARY(16) PRIMARY KEY,
    name VARCHAR(100),
    created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);

-- Converting a UUID to BINARY for storage
INSERT INTO users (id, name) 
VALUES (UNHEX(REPLACE(UUID(), '-', '')), 'John Doe');

-- Reading it back
SELECT HEX(id), name FROM users;

For MySQL, the recommendation is:

  1. Store UUIDs as BINARY(16) (saves memory)
  2. Use UUIDv7 for sortability
  3. Convert with UNHEX() on insert

ULID as an alternative to UUID

ULID (Universally Unique Lexicographically Sortable Identifier) is a UUID alternative built specifically for databases.

01ARZ3NDEKTSV4RRFFQ69G5FAV

Advantages of ULID:

  • Lexicographically sortable (timestamp at the front)
  • Shorter as a string (26 characters vs. 36 for UUID)
  • Base32 encoding (URL-safe)
  • Monotonic within a millisecond

Format:

  • 48 bits of timestamp (milliseconds)
  • 80 bits of randomness

ULID is essentially the same idea as UUIDv7, just with different encoding. For new projects, either one works fine.

// Go
import "github.com/oklog/ulid/v2"

entropy := rand.New(rand.NewSource(time.Now().UnixNano()))
id := ulid.MustNew(ulid.Timestamp(time.Now()), entropy)
fmt.Println(id)  // 01ARZ3NDEKTSV4RRFFQ69G5FAV
# Python
from ulid import ULID

ulid = ULID()
print(ulid)  # 01ARZ3NDEKTSV4RRFFQ69G5FAV

Best Practices

For databases:

  1. Use UUIDv7 (or ULID) for primary keys — the best performance
  2. Don’t use UUIDv4 for primary keys in high-load systems
  3. In PostgreSQL, use the UUID type, not VARCHAR(36)
  4. In MySQL, store it as BINARY(16)
  5. Create indexes on frequently queried fields

For APIs:

  1. Use UUIDv4 for public IDs — for security
  2. Don’t expose internal IDs (auto-increment) through the API
  3. Validate UUIDs on input (format and version)
  4. Use name-based UUIDs (v5) for idempotency

For distributed systems:

  1. Use UUIDv7 — it’s sortable and secure
  2. Generate on the client side (offline-first)
  3. Don’t rely on the timestamp for business logic (it may not be accurate)
  4. Keep server clocks in sync (NTP)

For security:

  1. Use UUIDv4 for tokens and sessions
  2. Don’t use v1/v6 in public APIs (they expose the MAC address)
  3. Don’t rely on the unpredictability of name-based UUIDs (v3/v5)
  4. Use a cryptographically secure random source for v4

Migrating from auto-increment to UUID

If you already have a system with auto-increment IDs and want to migrate to UUID, here’s a strategy:

Option 1: Add a new column

-- Add a UUID column
ALTER TABLE users ADD COLUMN uuid UUID;

-- Generate UUIDs for existing rows
UPDATE users SET uuid = gen_random_uuid() WHERE uuid IS NULL;

-- Make it NOT NULL
ALTER TABLE users ALTER COLUMN uuid SET NOT NULL;

-- Create a unique index
CREATE UNIQUE INDEX idx_users_uuid ON users(uuid);

-- Gradually switch the application over to using the UUID
-- Then you can drop the old id column

Option 2: Name-based UUID (preserve determinism)

import uuid

def migrate_id_to_uuid(old_id: int) -> uuid.UUID:
    namespace = uuid.UUID('6ba7b810-9dad-11d1-80b4-00c04fd430c8')
    return uuid.uuid5(namespace, f"user:{old_id}")

# In SQL
UPDATE users SET uuid = uuid_generate_v5(
    '6ba7b810-9dad-11d1-80b4-00c04fd430c8'::uuid,
    'user:' || id::text
);

This way the old ID can always be converted back into the same UUID.

Practical usage examples

Example 1: Microservices architecture (Go)

package main

import (
	"fmt"
	"time"
	"github.com/gofrs/uuid"
)

type Order struct {
	ID        uuid.UUID
	UserID    uuid.UUID
	CreatedAt time.Time
}

func NewOrder(userID uuid.UUID) (*Order, error) {
	// Generate a UUIDv7 for optimal database performance
	orderID, err := uuid.NewV7()
	if err != nil {
		return nil, err
	}

	return &Order{
		ID:        orderID,
		UserID:    userID,
		CreatedAt: time.Now(),
	}, nil
}

func main() {
	// UUIDv7 sorts automatically by creation time
	order1, _ := NewOrder(uuid.Must(uuid.NewV4()))
	time.Sleep(10 * time.Millisecond)
	order2, _ := NewOrder(uuid.Must(uuid.NewV4()))

	fmt.Println(order1.ID.String())  // 018d3f51-8b00-7000-9000-123456789abc
	fmt.Println(order2.ID.String())  // 018d3f51-8b20-7000-9000-234567890bcd
	// Sorted lexicographically by time
}

Example 2: Content deduplication (Python)

import uuid

class ContentDeduplicator:
    def __init__(self):
        self.namespace = uuid.UUID('6ba7b810-9dad-11d1-80b4-00c04fd430c8')
    
    def get_content_id(self, content: str) -> uuid.UUID:
        # Generate a deterministic UUID from the content
        return uuid.uuid5(self.namespace, content)
    
    def is_duplicate(self, content: str, seen_ids: set) -> bool:
        content_id = self.get_content_id(content)
        if content_id in seen_ids:
            return True
        seen_ids.add(content_id)
        return False

# Usage
dedup = ContentDeduplicator()
seen = set()

articles = [
    "Hello, world!",
    "Hello, world!",  # Duplicate
    "Different content"
]

for article in articles:
    if dedup.is_duplicate(article, seen):
        print(f"Duplicate: {article[:20]}...")
    else:
        print(f"New: {article[:20]}...")

Example 3: Idempotent API (JavaScript/TypeScript)

import { v5 as uuidv5, v4 as uuidv4 } from 'uuid';

class OrderService {
  private namespace = '6ba7b810-9dad-11d1-80b4-00c04fd430c8';

  // Idempotent order creation
  async createOrder(userId: string, items: any[], idempotencyKey?: string) {
    let orderId: string;

    if (idempotencyKey) {
      // Use a name-based UUID for idempotency
      orderId = uuidv5(idempotencyKey, this.namespace);
      
      // Check whether the order already exists
      const existing = await this.findOrderById(orderId);
      if (existing) {
        return existing;  // Return the existing one
      }
    } else {
      // Generate a new random UUID
      orderId = uuidv4();
    }

    // Create the order
    return await this.saveOrder({ id: orderId, userId, items });
  }

  private async findOrderById(id: string) {
    // DB lookup logic
  }

  private async saveOrder(order: any) {
    // Save logic
  }
}

// Usage
const service = new OrderService();

// With an idempotency key — always creates just one order
await service.createOrder('user123', [item1, item2], 'request-abc-123');
await service.createOrder('user123', [item1, item2], 'request-abc-123');  // Returns the same one

// Without an idempotency key — creates a new order every time
await service.createOrder('user123', [item1, item2]);
await service.createOrder('user123', [item1, item2]);  // Creates a second one

Conclusion

UUID is a powerful tool for distributed systems and microservices architectures. But the version you pick matters:

For databases:

  • Use UUIDv7 (timestamp-based, sortable) — the best performance
  • Avoid UUIDv4 for primary keys in high-load systems

For APIs:

  • Use UUIDv4 (random) — for security and unpredictability
  • Use UUIDv5 (name-based) for idempotency and caching

For deduplication:

  • Use UUIDv5 (name-based) — for determinism

For migration:

  • Use UUIDv5 (name-based) — you can reconstruct it from the old IDs
  • Or add a new column with UUIDv7 for optimal performance

Start with UUIDv7 for databases and UUIDv4 for everything else — that covers 90% of cases. And if you need something specific (deduplication, idempotency), reach for name-based UUIDv5.

If you’re getting into microservices, check out my article on adopting gRPC in a Go project — it covers how to organize communication between services and share proto files. And for automating routine tasks, I recommend Cursor AI — it’s great at helping with code generation.


P.S. Don’t use auto-increment IDs in public APIs — it’s trivial to enumerate all your records. Use UUIDs instead. And don’t forget indexes on your UUID columns.