Concurrency in Go: Channels vs Sync — When to Use Which

A deep dive into concurrency in Go: when to use channels and when to use sync.Mutex. Performance comparison, worker pool and fan-out/fan-in patterns, correct usage of WaitGroup, Once, and Pool. Practical examples and benchmarks.
- tags
- #Go #Programming #Concurrency #Performance
- categories
- Programming
- published
Working with Go? Sooner or later you’ll run into concurrency. Go promotes the philosophy “Don’t communicate by sharing memory, share memory by communicating” — and everyone immediately rushes to use channels everywhere. But that’s not always the right call.
I recently profiled a Go service and found that replacing channels with a mutex in the critical paths gave a 3x performance boost. Let’s go over when to use what, and why blindly following the “Go way” can tank your performance.
Concurrency vs Parallelism
First, let’s clear up the terminology, since a lot of people mix these up.
Concurrency is a program structure that lets you run multiple tasks by switching between them. It can work on a single core.
Parallelism is running multiple tasks at the same time on different cores.
// Concurrency — 2 goroutines on 1 core (switching between them)
runtime.GOMAXPROCS(1)
go task1()
go task2()
// Parallelism — 2 goroutines on 2 cores (running simultaneously)
runtime.GOMAXPROCS(2)
go task1()
go task2()
Go gives you concurrency out of the box via goroutines. Parallelism depends on GOMAXPROCS (which defaults to the number of cores).
Channels: the basics
A channel is a typed pipe for passing data between goroutines. It’s the core concurrency primitive in Go.
// Unbuffered channel — blocks the sender until someone receives
ch := make(chan int)
// Buffered channel — blocks only once the buffer is full
ch := make(chan int, 10)
// Send
ch <- 42
// Receive
value := <-ch
// Close
close(ch)
Unbuffered vs buffered channels
Unbuffered channel — synchronization. The sender blocks until the receiver reads the value.
func main() {
ch := make(chan int) // unbuffered
go func() {
fmt.Println("Sending...")
ch <- 42 // Blocks here
fmt.Println("Sent!")
}()
time.Sleep(time.Second)
fmt.Println("Receiving...")
value := <-ch // Unblocks the sender
fmt.Println("Received:", value)
}
// Output:
// Sending...
// (1 second pause)
// Receiving...
// Sent!
// Received: 42
Buffered channel — asynchronous until the buffer fills up.
func main() {
ch := make(chan int, 2) // buffer for 2 elements
ch <- 1 // Doesn't block
ch <- 2 // Doesn't block
// ch <- 3 // Would block — buffer is full
fmt.Println(<-ch) // 1
fmt.Println(<-ch) // 2
}
Select — multiplexing channels
select lets you wait on multiple channels at once:
func worker(done <-chan struct{}, tasks <-chan Task) {
for {
select {
case <-done:
fmt.Println("Worker stopping")
return
case task := <-tasks:
process(task)
}
}
}
Non-blocking operations with default:
select {
case msg := <-ch:
fmt.Println("Received:", msg)
default:
fmt.Println("No message available")
}
Timeout:
select {
case result := <-ch:
return result, nil
case <-time.After(5 * time.Second):
return nil, errors.New("timeout")
}
The sync package: the basics
The sync package provides low-level synchronization primitives.
sync.Mutex — mutual exclusion
Protects shared state from concurrent access:
type Counter struct {
mu sync.Mutex
value int
}
func (c *Counter) Inc() {
c.mu.Lock()
defer c.mu.Unlock()
c.value++
}
func (c *Counter) Value() int {
c.mu.Lock()
defer c.mu.Unlock()
return c.value
}
sync.RWMutex — readers and writers
When you have many readers and few writers, RWMutex is more efficient:
type Cache struct {
mu sync.RWMutex
data map[string]string
}
func (c *Cache) Get(key string) (string, bool) {
c.mu.RLock() // Multiple readers can run concurrently
defer c.mu.RUnlock()
val, ok := c.data[key]
return val, ok
}
func (c *Cache) Set(key, value string) {
c.mu.Lock() // Only one writer at a time
defer c.mu.Unlock()
c.data[key] = value
}
sync.WaitGroup — waiting for completion
Waits for a group of goroutines to finish:
func processItems(items []Item) {
var wg sync.WaitGroup
for _, item := range items {
wg.Add(1)
go func(item Item) {
defer wg.Done()
process(item)
}(item)
}
wg.Wait() // Wait for all of them
}
Important: wg.Add() must be called BEFORE launching the goroutine, not inside it!
// WRONG — race condition
for _, item := range items {
go func(item Item) {
wg.Add(1) // Might not run before wg.Wait()
defer wg.Done()
process(item)
}(item)
}
wg.Wait()
// CORRECT
for _, item := range items {
wg.Add(1)
go func(item Item) {
defer wg.Done()
process(item)
}(item)
}
wg.Wait()
sync.Once — run exactly once
Guarantees a piece of code runs exactly once (initialization, singletons):
type Database struct {
once sync.Once
conn *sql.DB
}
func (d *Database) Connection() *sql.DB {
d.once.Do(func() {
// Runs only once, even under concurrent calls
d.conn, _ = sql.Open("postgres", "...")
})
return d.conn
}
sync.Pool — object reuse
Reduces GC pressure for objects that get created frequently:
var bufferPool = sync.Pool{
New: func() interface{} {
return new(bytes.Buffer)
},
}
func process(data []byte) {
buf := bufferPool.Get().(*bytes.Buffer)
defer func() {
buf.Reset()
bufferPool.Put(buf)
}()
buf.Write(data)
// work with buf
}
sync.Map — a concurrent map
Useful when keys are stable or there are far more readers than writers:
var cache sync.Map
// Write
cache.Store("key", "value")
// Read
if val, ok := cache.Load("key"); ok {
fmt.Println(val.(string))
}
// LoadOrStore — atomic
actual, loaded := cache.LoadOrStore("key", "default")
// Delete
cache.Delete("key")
// Iterate
cache.Range(func(key, value interface{}) bool {
fmt.Println(key, value)
return true // continue iterating
})
Important: sync.Map is NOT always faster than map + RWMutex. Use it when:
- Keys are stable (new ones are rarely added)
- There are many readers and few writers
- Goroutines are working with different keys
sync.Cond — condition variables
For complex coordination (rarely needed):
type Queue struct {
mu sync.Mutex
cond *sync.Cond
items []int
}
func NewQueue() *Queue {
q := &Queue{}
q.cond = sync.NewCond(&q.mu)
return q
}
func (q *Queue) Push(item int) {
q.mu.Lock()
q.items = append(q.items, item)
q.cond.Signal() // Wake up one waiter
q.mu.Unlock()
}
func (q *Queue) Pop() int {
q.mu.Lock()
for len(q.items) == 0 {
q.cond.Wait() // Wait for a signal
}
item := q.items[0]
q.items = q.items[1:]
q.mu.Unlock()
return item
}
When to use channels
Channels are a good fit for:
1. Transferring ownership of data
When one goroutine “hands off” data to another and stops using it:
func producer(out chan<- Task) {
for {
task := createTask()
out <- task // Transfer ownership
// No longer touch task
}
}
func consumer(in <-chan Task) {
for task := range in {
process(task) // We now own task
}
}
2. Coordination and signaling
Graceful shutdown, cancelling operations:
func server(done <-chan struct{}) {
for {
select {
case <-done:
fmt.Println("Shutting down...")
return
default:
handleRequest()
}
}
}
func main() {
done := make(chan struct{})
go server(done)
// ... do work ...
close(done) // Signal shutdown
}
3. The pipeline pattern
A chain of data-processing stages:
func gen(nums ...int) <-chan int {
out := make(chan int)
go func() {
for _, n := range nums {
out <- n
}
close(out)
}()
return out
}
func sq(in <-chan int) <-chan int {
out := make(chan int)
go func() {
for n := range in {
out <- n * n
}
close(out)
}()
return out
}
func main() {
// Pipeline: gen -> sq -> print
for n := range sq(gen(1, 2, 3, 4)) {
fmt.Println(n) // 1, 4, 9, 16
}
}
4. Fan-out / fan-in
Distributing work and collecting results:
// Fan-out: one source, many processors
func fanOut(in <-chan Task, workers int) []<-chan Result {
outs := make([]<-chan Result, workers)
for i := 0; i < workers; i++ {
outs[i] = worker(in)
}
return outs
}
// Fan-in: many sources, one receiver
func fanIn(channels ...<-chan Result) <-chan Result {
var wg sync.WaitGroup
out := make(chan Result)
for _, ch := range channels {
wg.Add(1)
go func(c <-chan Result) {
defer wg.Done()
for v := range c {
out <- v
}
}(ch)
}
go func() {
wg.Wait()
close(out)
}()
return out
}
5. Rate limiting
Limiting the rate of processing:
func rateLimited(requests <-chan Request) {
// 10 requests per second
limiter := time.NewTicker(100 * time.Millisecond)
defer limiter.Stop()
for req := range requests {
<-limiter.C // Wait for a tick
go handle(req)
}
}
// Burst rate limiting
func burstRateLimited(requests <-chan Request) {
// Allow a burst of up to 3, then 1 per second
burstyLimiter := make(chan time.Time, 3)
// Fill the burst
for i := 0; i < 3; i++ {
burstyLimiter <- time.Now()
}
// Refill at 1/sec
go func() {
for t := range time.Tick(time.Second) {
burstyLimiter <- t
}
}()
for req := range requests {
<-burstyLimiter
go handle(req)
}
}
When to use sync
Mutexes and the other sync primitives are better for:
1. Protecting shared state
When you just need to protect a variable:
// With a mutex — simple and fast
type Stats struct {
mu sync.Mutex
requests int
errors int
}
func (s *Stats) RecordRequest() {
s.mu.Lock()
s.requests++
s.mu.Unlock()
}
func (s *Stats) RecordError() {
s.mu.Lock()
s.errors++
s.mu.Unlock()
}
// With channels — more complex and slower
type Stats struct {
requests chan int
errors chan int
data struct {
requests int
errors int
}
}
func NewStats() *Stats {
s := &Stats{
requests: make(chan int),
errors: make(chan int),
}
go s.run()
return s
}
func (s *Stats) run() {
for {
select {
case <-s.requests:
s.data.requests++
case <-s.errors:
s.data.errors++
}
}
}
2. Read-heavy workloads
When there are many readers and few writers — use RWMutex:
type ConfigStore struct {
mu sync.RWMutex
config Config
}
func (c *ConfigStore) Get() Config {
c.mu.RLock()
defer c.mu.RUnlock()
return c.config
}
func (c *ConfigStore) Update(cfg Config) {
c.mu.Lock()
defer c.mu.Unlock()
c.config = cfg
}
3. Initialization (sync.Once)
Lazy singleton initialization:
var (
instance *Database
once sync.Once
)
func GetDatabase() *Database {
once.Do(func() {
instance = &Database{}
instance.connect()
})
return instance
}
4. Object reuse (sync.Pool)
Reducing allocations:
var jsonPool = sync.Pool{
New: func() interface{} {
return &bytes.Buffer{}
},
}
func MarshalJSON(v interface{}) ([]byte, error) {
buf := jsonPool.Get().(*bytes.Buffer)
defer func() {
buf.Reset()
jsonPool.Put(buf)
}()
enc := json.NewEncoder(buf)
if err := enc.Encode(v); err != nil {
return nil, err
}
return buf.Bytes(), nil
}
Benchmarks: channels vs mutex
Let’s measure the performance. The task: increment a counter from multiple goroutines.
// benchmark_test.go
package main
import (
"sync"
"sync/atomic"
"testing"
)
// Option 1: Mutex
type MutexCounter struct {
mu sync.Mutex
value int64
}
func (c *MutexCounter) Inc() {
c.mu.Lock()
c.value++
c.mu.Unlock()
}
// Option 2: Channel
type ChannelCounter struct {
inc chan struct{}
value int64
}
func NewChannelCounter() *ChannelCounter {
c := &ChannelCounter{
inc: make(chan struct{}),
}
go c.run()
return c
}
func (c *ChannelCounter) run() {
for range c.inc {
c.value++
}
}
func (c *ChannelCounter) Inc() {
c.inc <- struct{}{}
}
// Option 3: Atomic
type AtomicCounter struct {
value int64
}
func (c *AtomicCounter) Inc() {
atomic.AddInt64(&c.value, 1)
}
// Benchmarks
func BenchmarkMutexCounter(b *testing.B) {
c := &MutexCounter{}
b.RunParallel(func(pb *testing.PB) {
for pb.Next() {
c.Inc()
}
})
}
func BenchmarkChannelCounter(b *testing.B) {
c := NewChannelCounter()
b.RunParallel(func(pb *testing.PB) {
for pb.Next() {
c.Inc()
}
})
}
func BenchmarkAtomicCounter(b *testing.B) {
c := &AtomicCounter{}
b.RunParallel(func(pb *testing.PB) {
for pb.Next() {
c.Inc()
}
})
}
Results (MacBook Pro M1, 8 cores):
BenchmarkMutexCounter-8 50000000 24.5 ns/op
BenchmarkChannelCounter-8 10000000 145.0 ns/op
BenchmarkAtomicCounter-8 200000000 6.2 ns/op
| Approach | ns/op | Relative |
|---|---|---|
| Atomic | 6.2 | 1x (baseline) |
| Mutex | 24.5 | 4x slower |
| Channel | 145.0 | 23x slower |
Takeaway: For simple operations on shared state, a mutex is 6x faster than a channel. Atomic is another 4x faster than a mutex.
When are channels faster?
Channels win when the coordination is complex enough that the alternative would be a pile of mutexes and condition variables:
// Pipeline with channels — simple and effective
func pipeline(in <-chan int) <-chan int {
out := make(chan int)
go func() {
for v := range in {
out <- v * 2
}
close(out)
}()
return out
}
// The same thing with a mutex — a nightmare
type PipelineStage struct {
mu sync.Mutex
cond *sync.Cond
input []int
output []int
closed bool
}
// ... 50+ lines of code ...
Concurrency patterns
Worker pool
A fixed number of workers process tasks from a queue:
func workerPool(numWorkers int, tasks <-chan Task, results chan<- Result) {
var wg sync.WaitGroup
for i := 0; i < numWorkers; i++ {
wg.Add(1)
go func(workerID int) {
defer wg.Done()
for task := range tasks {
result := process(task)
results <- result
}
}(i)
}
wg.Wait()
close(results)
}
func main() {
tasks := make(chan Task, 100)
results := make(chan Result, 100)
// Start 10 workers
go workerPool(10, tasks, results)
// Send tasks
go func() {
for _, t := range getAllTasks() {
tasks <- t
}
close(tasks)
}()
// Collect results
for result := range results {
handleResult(result)
}
}
Semaphore (limiting concurrency)
Limiting how many operations run at the same time:
// Semaphore built on a buffered channel
type Semaphore chan struct{}
func NewSemaphore(n int) Semaphore {
return make(chan struct{}, n)
}
func (s Semaphore) Acquire() {
s <- struct{}{}
}
func (s Semaphore) Release() {
<-s
}
// Usage
func processWithLimit(items []Item, maxConcurrency int) {
sem := NewSemaphore(maxConcurrency)
var wg sync.WaitGroup
for _, item := range items {
wg.Add(1)
sem.Acquire()
go func(item Item) {
defer wg.Done()
defer sem.Release()
process(item)
}(item)
}
wg.Wait()
}
Context for cancellation
Use context for graceful cancellation:
func worker(ctx context.Context, tasks <-chan Task) error {
for {
select {
case <-ctx.Done():
return ctx.Err()
case task, ok := <-tasks:
if !ok {
return nil // Channel closed
}
if err := process(ctx, task); err != nil {
return err
}
}
}
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
tasks := make(chan Task)
// Start the workers
g, ctx := errgroup.WithContext(ctx)
for i := 0; i < 10; i++ {
g.Go(func() error {
return worker(ctx, tasks)
})
}
// Send tasks
go func() {
defer close(tasks)
for _, t := range getAllTasks() {
select {
case <-ctx.Done():
return
case tasks <- t:
}
}
}()
if err := g.Wait(); err != nil {
log.Fatal(err)
}
}
errgroup — coordinating with errors
The golang.org/x/sync/errgroup package makes working with a group of goroutines much easier:
import "golang.org/x/sync/errgroup"
func fetchAll(ctx context.Context, urls []string) ([]Response, error) {
g, ctx := errgroup.WithContext(ctx)
results := make([]Response, len(urls))
for i, url := range urls {
i, url := i, url // Capture
g.Go(func() error {
resp, err := fetch(ctx, url)
if err != nil {
return err
}
results[i] = resp
return nil
})
}
if err := g.Wait(); err != nil {
return nil, err
}
return results, nil
}
Singleflight — deduplicating requests
Preventing a thundering herd (many identical requests at once):
import "golang.org/x/sync/singleflight"
var requestGroup singleflight.Group
func getData(key string) ([]byte, error) {
// All concurrent requests for the same key will trigger only one fetch
v, err, _ := requestGroup.Do(key, func() (interface{}, error) {
return fetchFromDB(key)
})
if err != nil {
return nil, err
}
return v.([]byte), nil
}
Very useful for caching: if 1000 requests come in at once for the same key, the database only sees 1 query.
Race conditions and how to avoid them
The race detector
Go has a built-in race condition detector:
go test -race ./...
go build -race ./cmd/server
Example of a race condition:
// RACE CONDITION
func main() {
counter := 0
for i := 0; i < 1000; i++ {
go func() {
counter++ // Unsynchronized access
}()
}
time.Sleep(time.Second)
fmt.Println(counter) // Undefined result
}
$ go run -race main.go
==================
WARNING: DATA RACE
Read at 0x00c0000a4010 by goroutine 8:
main.main.func1()
/main.go:11 +0x3c
Previous write at 0x00c0000a4010 by goroutine 7:
main.main.func1()
/main.go:11 +0x52
==================
Fixing it with a mutex
func main() {
var mu sync.Mutex
counter := 0
for i := 0; i < 1000; i++ {
go func() {
mu.Lock()
counter++
mu.Unlock()
}()
}
time.Sleep(time.Second)
fmt.Println(counter) // 1000
}
Fixing it with atomic
func main() {
var counter int64
for i := 0; i < 1000; i++ {
go func() {
atomic.AddInt64(&counter, 1)
}()
}
time.Sleep(time.Second)
fmt.Println(counter) // 1000
}
Fixing it with a channel
func main() {
counter := 0
inc := make(chan struct{}, 1000)
for i := 0; i < 1000; i++ {
go func() {
inc <- struct{}{}
}()
}
for i := 0; i < 1000; i++ {
<-inc
counter++
}
fmt.Println(counter) // 1000
}
Common mistakes
1. Goroutine leaks
// LEAK — the goroutine will never finish
func leaky() {
ch := make(chan int)
go func() {
val := <-ch // Blocks forever
fmt.Println(val)
}()
// ch will never receive a value or get closed
}
// FIX — use context for cancellation
func notLeaky(ctx context.Context) {
ch := make(chan int)
go func() {
select {
case val := <-ch:
fmt.Println(val)
case <-ctx.Done():
return
}
}()
}
2. Closing a channel more than once
// PANIC
ch := make(chan int)
close(ch)
close(ch) // panic: close of closed channel
// FIX — close it only once
type SafeChannel struct {
ch chan int
once sync.Once
closed bool
}
func (s *SafeChannel) Close() {
s.once.Do(func() {
close(s.ch)
s.closed = true
})
}
3. Sending on a closed channel
// PANIC
ch := make(chan int)
close(ch)
ch <- 1 // panic: send on closed channel
// Only the sender should ever close the channel!
func producer(ch chan<- int) {
defer close(ch) // Close once we're done sending
for i := 0; i < 10; i++ {
ch <- i
}
}
4. Copying a mutex
// WRONG — the mutex gets copied
type Counter struct {
mu sync.Mutex
value int
}
func (c Counter) Inc() { // Value receiver — copies the mutex!
c.mu.Lock()
c.value++
c.mu.Unlock()
}
// CORRECT — pointer receiver
func (c *Counter) Inc() {
c.mu.Lock()
c.value++
c.mu.Unlock()
}
5. WaitGroup inside a goroutine
// WRONG — race condition
var wg sync.WaitGroup
for i := 0; i < 10; i++ {
go func() {
wg.Add(1) // Might not run before wg.Wait()
defer wg.Done()
work()
}()
}
wg.Wait()
// CORRECT — Add before launching the goroutine
var wg sync.WaitGroup
for i := 0; i < 10; i++ {
wg.Add(1)
go func() {
defer wg.Done()
work()
}()
}
wg.Wait()
6. Capturing the loop variable
// WRONG — all goroutines get the final value of i
for i := 0; i < 10; i++ {
go func() {
fmt.Println(i) // All of them print 10
}()
}
// CORRECT — pass the value explicitly
for i := 0; i < 10; i++ {
go func(i int) {
fmt.Println(i) // 0, 1, 2, ..., 9 (in random order)
}(i)
}
// Or a local variable
for i := 0; i < 10; i++ {
i := i // Shadow
go func() {
fmt.Println(i)
}()
}
Note: In Go 1.22+ this issue is fixed — each iteration gets its own fresh variable.
Table: when to use what
| Task | Solution | Why |
|---|---|---|
| Protecting a variable | sync.Mutex | Simple and fast |
| Many readers, few writers | sync.RWMutex | Concurrent reads |
| Counter | atomic | The fastest option |
| One-time initialization | sync.Once | Thread-safe singleton |
| Passing data between goroutines | channel | Ownership transfer |
| Graceful shutdown | channel or context | Signaling |
| Processing pipeline | channel | A natural fit |
| Worker pool | channel + sync.WaitGroup | Coordinating workers |
| Limiting concurrency | Buffered channel (semaphore) | Simple semaphore |
| Object reuse | sync.Pool | Reducing allocations |
| Concurrent map | sync.Map or map + RWMutex | Depends on the access pattern |
| Deduplicating requests | singleflight | Preventing thundering herd |
| A group of goroutines with error handling | errgroup | Convenient coordination |
Best practices
1. Start simple
// Start with a mutex — simple and clear
type Cache struct {
mu sync.RWMutex
data map[string]string
}
// Move to channels only if you need complex coordination
2. Use -race in tests and CI
# .github/workflows/test.yml
- name: Test with race detector
run: go test -race -v ./...
3. Document your synchronization rules
// Counter is safe for concurrent use.
// All methods may be called from multiple goroutines.
type Counter struct {
mu sync.Mutex
value int
}
4. Keep the critical section short
// BAD — a slow operation held under the lock
func (c *Cache) Process(key string) {
c.mu.Lock()
defer c.mu.Unlock()
data := c.data[key]
result := expensiveOperation(data) // Slow!
c.data[key] = result
}
// GOOD — the lock only covers data access
func (c *Cache) Process(key string) {
c.mu.RLock()
data := c.data[key]
c.mu.RUnlock()
result := expensiveOperation(data) // Outside the lock
c.mu.Lock()
c.data[key] = result
c.mu.Unlock()
}
5. Use context for cancellation
func worker(ctx context.Context) error {
for {
select {
case <-ctx.Done():
return ctx.Err()
default:
if err := doWork(ctx); err != nil {
return err
}
}
}
}
6. Close channels on the sender’s side
func producer(out chan<- int) {
defer close(out) // The sender closes it
for i := 0; i < 10; i++ {
out <- i
}
}
func consumer(in <-chan int) {
for v := range in { // The receiver ranges until it's closed
fmt.Println(v)
}
}
Conclusion
Concurrency in Go is a powerful tool, but it’s important to pick the right primitive:
Use channels when:
- You’re transferring ownership of data
- You need coordination and signaling
- You’re building a pipeline
- You’re implementing graceful shutdown
Use sync when:
- You’re protecting shared state
- You need maximum performance
- There are many readers and few writers
- You’re initializing a singleton
Rule of thumb:
When in doubt, start with a mutex. Move to channels only when the code becomes more complex than it would be with channels.
Remember: “Don’t communicate by sharing memory” is a guideline, not a law. Sometimes a mutex is the right choice.
If you want to dig deeper into Go, check out the article on Go anti-patterns — it has real examples of bad code from production. And for working with microservices, there’s an article on gRPC in Go .
P.S. Always run your tests with -race. It’s better to catch a race condition in CI than in production at 3am.