Skip to content
All projects

Project · 03

Distributed Binary Data Sorter

Distributed Computing / Data Processing

Sorts a big list of numbers across multiple machines. Each one sorts its own chunk, everything stays in binary along the way, and the results get merged back together as they come in.

Year
2026
Stack
Python · Dispy · struct · threading

Overview

One machine (the coordinator) breaks a big file of numbers into chunks and sends them out to worker machines with Dispy. Each worker sorts its chunk, and the coordinator merges finished chunks two at a time while the rest are still running.

In the final version everything stays in binary — numbers get packed into 4-byte ints before they're sent out and stay that way through every merge, until the very end when it's written out as text.

  • Streams the input and packs chunks into 4-byte integers
  • Submits each chunk to worker nodes as soon as it's ready
  • Workers sort chunks read from shared storage
  • Sorted results written back in binary
  • Pairwise merging as jobs complete
  • Lock-protected merge queue across job callbacks
  • Per-job error reporting
  • Final merge converts to text output

Focus

  • Distributed computing
  • Worker nodes
  • Binary file I/O
  • Synchronization
  • Incremental merging

Architecture

The coordinator reads the input as it goes instead of loading it all at once, packs each chunk with Python's struct module, and sends it off right away — so workers start sorting before the whole file has even been read. Every time a job finishes, a callback adds its output to a queue and merges pairs of files, with a lock so callbacks don't step on each other.

Source

  • Input dataset
  1. 01Chunk + pack to binary
  2. 02Dispatch jobs (Dispy)
  3. 03Sort on worker nodes
  4. 04Incremental binary merge
  5. 05Final text output
Sorting pipeline: Input dataset → Chunk + pack to binary → Dispatch jobs (Dispy) → Sort on worker nodes → Incremental binary merge → Final text output

Evolution

I built this in three steps: get a distributed sort working, figure out where it was doing extra work, test a fix on its own, then put it back in.

  1. Multinode Data Sorter

    • Split a big text file into chunks on shared storage
    • Sent the sorting work out to multiple worker machines
    • Merged the sorted files back on the coordinator — but every merge had to read and rewrite each number as text
  2. Binary Merge Utility

    • Merged two sorted files of numbers straight in binary
    • Compared the raw 4-byte values instead of reading text lines
  3. Binary-Optimized Multinode Data Sorter

    • Turn chunks into binary
    • Send binary chunks to workers as soon as they're ready
    • Sort on the workers
    • Merge the binary results as each job finishes
    • Write the final result out as text

Stack

  • Python
  • Dispy
  • struct
  • threading

Contact

Get in touch.

I'm looking for software engineering internships. Email's the fastest way to reach me.