I'm building a Python project that needs to save and load variables containing many different kinds of data, from simple integers and strings to dictionaries, custom objects, and potentially large NumPy arrays. I initially used pickle, but its output is not deterministic enough for my needs. I tried creating a custom native C++ serializer that handles built-in types directly, caches class layouts, batches similar objects, bulk-copies homogeneous arrays, and skips pickle's memo table. However, it is almost twice as slow as pickle. Disk size is not important; serialization speed is the main concern. Is there a faster approach or library that can provide deterministic serialization for varied Python objects?
3 Answers
The best option depends heavily on the actual data. If most values are numeric arrays, strings, or fixed layouts, a format built around raw buffers—such as NumPy’s binary formats, `struct`, or `array`—can be very fast. For arbitrary mixtures of dictionaries, custom objects, and arrays, though, you’ll usually need either a schema or some conversion layer, so there may not be a universally fast drop-in replacement.
Consider benchmarking `msgspec`’s MessagePack encoder with deterministic ordering enabled. It can be very fast, but sorting keys adds overhead and custom classes may require conversion hooks, so include that work in the benchmark. Also, removing pickle’s memo table is not automatically an optimization: if the object graph contains repeated references, duplicating those objects can make serialization slower. Pickle is already implemented largely in optimized C, so moving parts of the implementation to C++ will not necessarily beat it. Profile the real workload before choosing a format.
It would help to define exactly what deterministic means here. Are you trying to produce identical bytes across separate runs, make dictionary ordering consistent, or keep the format stable across Python versions? Also benchmark the current implementation with concrete object counts, sizes, and timings. Without those numbers, it’s difficult to tell whether serialization, object traversal, or file I/O is the bottleneck.

The data is fairly general: simple values, dictionaries, custom objects, and sometimes complex NumPy data. I’m looking for something that can handle varied types without requiring a completely fixed schema.