Bug report
Bug description:
The GNU and ustar tar format families have slightly overlapping header definitions: ustar says that bytes 345:500 are used for the "path prefix" i.e. for paths longer than the 100 char limit in the path field, while old-style GNU ("oldgnu") says that 345:356 are used for storing an atime attribute for the member.
Consequently, the two overlap (partially), and a tar parser should check the header's magic to determine how to interpret that byte range.
At the moment, CPython does not use the header's magic, and instead unconditionally interprets that range as a ustar-style prefix:
|
prefix = nts(buf[345:500], encoding, errors) |
and then unconditionally uses that prefix as long as the member type (not the header type) isn't a special GNU member type:
|
# Reconstruct a ustar longname. |
|
if prefix and obj.type not in GNU_TYPES: |
|
obj.name = prefix + "/" + obj.name |
The end result of this is that tarfile can extract a file with a surprising name, whereas other parsers extract with the correct (non-ustar-prefixed) name.
MRE:
import io
import tarfile
member = tarfile.TarInfo("victim")
header = bytearray(member.tobuf(format=tarfile.GNU_FORMAT))
# Old-GNU atime field: valid octal timestamp 1.
header[345:357] = b"00000000001\0"
# Recalculate checksum.
header[148:156] = b" " * 8
header[148:156] = f"{sum(header):06o}\0 ".encode("ascii")
archive = bytes(header) + b"\0" * 1024
with tarfile.open(fileobj=io.BytesIO(archive), mode="r:") as tf:
print(tf.getnames())
On a main build as of b11e749, this produces:
whereas the output should be ['victim'], since the format is GNU_FORMAT instead of a ustar-family format.
I think the fix for this is to tweak the obj.name assignment to only use prefix when the magic bytes match POSIX_MAGIC, i.e. not GNU_MAGIC or any legacy (v7, pre-ustar) magic.
CPython versions tested on:
CPython main branch
Operating systems tested on:
No response
Linked PRs
Bug report
Bug description:
The GNU and ustar tar format families have slightly overlapping header definitions: ustar says that bytes
345:500are used for the "path prefix" i.e. for paths longer than the 100 char limit in the path field, while old-style GNU ("oldgnu") says that345:356are used for storing anatimeattribute for the member.Consequently, the two overlap (partially), and a tar parser should check the header's magic to determine how to interpret that byte range.
At the moment, CPython does not use the header's magic, and instead unconditionally interprets that range as a ustar-style prefix:
cpython/Lib/tarfile.py
Line 1354 in b11e749
and then unconditionally uses that prefix as long as the member type (not the header type) isn't a special GNU member type:
cpython/Lib/tarfile.py
Lines 1383 to 1385 in b11e749
The end result of this is that
tarfilecan extract a file with a surprising name, whereas other parsers extract with the correct (non-ustar-prefixed) name.MRE:
On a main build as of b11e749, this produces:
whereas the output should be
['victim'], since the format isGNU_FORMATinstead of a ustar-family format.I think the fix for this is to tweak the
obj.nameassignment to only useprefixwhen the magic bytes matchPOSIX_MAGIC, i.e. notGNU_MAGICor any legacy (v7, pre-ustar) magic.CPython versions tested on:
CPython main branch
Operating systems tested on:
No response
Linked PRs