Insecure Deserialization & Integrity Failures

Deserialization turns bytes back into objects. When an application deserializes untrusted data with a format that can instantiate arbitrary types and run their code (Java's native serialization, Python's pickle, PHP's unserialize, .NET's BinaryFormatter, unsafe YAML loaders), an attacker can craft a payload that executes commands on the server the moment it's parsed. Insecure deserialization has produced some of the most severe vulnerabilities in widely used software, often with full remote code execution (RCE).

The 2021 OWASP Top 10 folds it into A08: Software and Data Integrity Failures, a broader category about trusting data or code without verifying its integrity. That includes unsigned updates, compromised CI/CD pipelines, dependencies from untrusted sources, and tampered cookies or tokens. The unifying lesson: never let untrusted input decide what code runs, and verify integrity before trusting data.

TL;DR

Quick Example

Python pickle is code execution by design:

Vulnerable pattern: storing a pickled object in a cookie or accepting uploaded pickles:

YAML:

Java deserialization filter (JEP 290) as defense in depth:

Core Concepts

Why Deserialization Can Execute Code

Native serialization formats encode type information along with data. During deserialization, the runtime creates instances of the named classes and invokes special methods: Java's readObject/readResolve, Python's __reduce__/__setstate__, PHP's __wakeup/__destruct, .NET serialization callbacks. Attackers chain together existing classes in the application's libraries (gadgets) whose methods, invoked in sequence during deserialization, end up executing commands, writing files, or making requests. Tools like ysoserial automate gadget chains for popular Java libraries, which is why "we don't have dangerous classes" is rarely true.

Dangerous APIs by Ecosystem

Where Untrusted Serialized Data Appears

Software and Data Integrity Failures

The broader A08 category also covers:

Supply chain controls include signing artifacts (Sigstore/cosign), verifying provenance (SLSA attestations), SBOMs, and pinned, hash-verified dependencies. See container security.

Best Practices

Use Data-Only Formats Across Trust Boundaries

Anything crossing a trust boundary (client ↔ server, service ↔ service, uploads, queues) should use formats that can't encode behavior: JSON, Protobuf, Avro, CBOR, or MessagePack, deserialized into explicit, validated types.

Authenticate Data You Must Round-Trip

If you store state in clients or other less-trusted locations (cookies, tokens, URLs), sign it with HMAC or digital signatures, and verify the signature before parsing, or use framework-provided signed or encrypted sessions. Signing doesn't make unsafe formats safe if the key leaks, so still prefer data-only formats.

Restrict Polymorphism

Disable polymorphic type handling driven by input (Jackson default typing, Json.NET TypeNameHandling), or constrain it to an allowlist of known subtypes, such as discriminated unions with fixed type names.

Treat Model Files as Code

Loading a pickled ML model is running code. Use safetensors or ONNX for weights, load models only from trusted, verified sources, and scan or sandbox third-party models. See MLOps.

Common Mistakes

"It's Internal, So It's Trusted"

Internal queues and caches become attack paths once any producer, network segment, or cache entry is compromised, including via SSRF into an unauthenticated Redis. Apply the same rules internally.

Blocklisting Known Gadget Classes

Blocklists of known-dangerous classes are bypassed by new gadget chains in newly added libraries. Use allowlists, or better, eliminate native deserialization of untrusted data.

Encrypting Instead of Authenticating

Encrypting a serialized blob without integrity protection (for example AES-CBC without a MAC) can still allow tampering and padding oracle attacks. Use authenticated encryption or signatures.

FAQ

What is insecure deserialization?

It's deserializing attacker-controlled data with a mechanism that can instantiate arbitrary objects and trigger their methods. Attackers craft payloads that abuse existing classes (gadget chains) to execute code, tamper with application logic, or cause denial of service when the application parses the data.

Is JSON deserialization safe?

Plain JSON parsing into primitive types or explicitly declared classes is generally safe from code execution. Risk returns when libraries let JSON specify which class to instantiate (polymorphic type handling driven by input), so keep that disabled or strictly allowlisted, and validate the resulting data.

Why is Python pickle dangerous?

Pickle is a small program that the unpickler executes. It can call any importable function with any arguments via __reduce__, so loading a malicious pickle is equivalent to running attacker code. Python's documentation explicitly warns never to unpickle untrusted data.

How does this relate to supply chain security?

Both are integrity failures: trusting data or code without verifying where it came from and that it wasn't modified. The defenses are similar: verify signatures and hashes, restrict trusted sources, and avoid mechanisms that execute whatever they're given.

Related Topics

References