I'm building a Python project that needs to save and load variables containing mixed data: integers, strings, dictionaries, custom objects, and potentially large NumPy arrays. I initially used pickle, but I need deterministic output and don't want serialization results to vary between runs or because of dictionary ordering.
I also tried writing a native C++ serializer that dispatches directly for built-in types, caches layouts for custom classes, batches homogeneous data, and uses memory-mapped loading. However, it is nearly twice as slow as pickle for my workload. Disk space is not important; my main goal is serialization and deserialization speed. Is there a faster approach or library that can provide deterministic output for this kind of varied data?
3 Answers
Benchmark msgspec’s MessagePack encoder with deterministic ordering enabled. It can be very fast, although sorting keys to produce stable output adds overhead and custom classes may require conversion hooks. Also define what deterministic means in your case: identical bytes across runs, canonical dictionary ordering, or compatibility across Python and library versions. Those requirements can significantly affect the design.
Pickle is already implemented largely in optimized C, so rewriting the serializer in C++ does not automatically make it faster. Its memo table is also useful when objects share references; removing it may cause large objects to be serialized repeatedly. Profile encoding and decoding separately, measure the size and shape of the data, and check whether conversion of custom objects or NumPy arrays dominates the runtime before changing formats.
The best option depends heavily on the actual object mix and benchmark numbers. If the data is mostly numeric arrays, strings, and simple values, a format based on raw buffers—such as struct, array, or a specialized NumPy format—can beat a general-purpose object serializer. For arbitrary nested Python objects and custom classes, though, you usually have to trade speed, portability, or supported object types.
The data is mixed: simple integers and strings as well as dictionaries, custom objects, and fairly complex NumPy data. The serializer needs to support a broad range of types rather than just homogeneous arrays.

If exact byte-for-byte output is required, you may need to canonicalize dictionaries and other unordered structures before serialization. That extra sorting work can make a deterministic serializer slower than ordinary pickle, even when the resulting format is compact.