Maven module :un.storage : archive-xz :
Class : un.storage.archive.xz.SeekableXZInputStream
Extends/Implements : un.storage.archive.xz.SeekableInputStream
Subclasses : -

Decompresses a .xz file in random access mode.
This supports decompressing concatenated .xz files.


Each .xz file consist of one or more Streams. Each Stream consist of zero
or more Blocks. Each Stream contains an Index of Streams' Blocks.
The Indexes from all Streams are loaded in RAM by a constructor of this
class. A typical .xz file has only one Stream, and parsing its Index will
need only three or four seeks.


To make random access possible, the data in a .xz file must be splitted
into multiple Blocks of reasonable size. Decompression can only start at
a Block boundary. When seeking to an uncompressed position that is not at
a Block boundary, decompression starts at the beginning of the Block and
throws away data until the target position is reached. Thus, smaller Blocks
mean faster seeks to arbitrary uncompressed positions. On the other hand,
smaller Blocks mean worse compression. So one has to make a compromise
between random access speed and compression ratio.


Implementation note: This class uses linear search to locate the correct
Stream from the data structures in RAM. It was the simplest to implement
and should be fine as long as there aren't too many Streams. The correct
Block inside a Stream is located using binary search and thus is fast
even with a huge number of Blocks.

Memory usage



The amount of memory needed for the Indexes is taken into account when
checking the memory usage limit. Each Stream is calculated to need at
least 1 KiB of memory and each Block 16 bytes of memory, rounded up
to the next kibibyte. So unless the file has a huge number of Streams or
Blocks, these don't take significant amount of memory.

Creating random-accessible .xz files



When using {@link XZOutputStream}, a new Block can be started by calling
its {@link XZOutputStream#endBlock() endBlock} method. If you know
that the decompressor will only need to seek to certain uncompressed
positions, it can be a good idea to start a new Block at (some of) these
positions (and only at these positions to get better compression ratio).


liblzma in XZ Utils supports starting a new Block with
LZMA_FULL_FLUSH. XZ Utils 5.1.1alpha added threaded
compression which creates multi-Block .xz files. XZ Utils 5.1.1alpha
also added the option --block-size=SIZE to the xz command
line tool. XZ Utils 5.1.2alpha added a partial implementation of
--block-list=SIZES which allows specifying sizes of
individual Blocks.

@see SeekableFileInputStream
@see XZInputStream
@see XZOutputStream



Variables : -
Functions : SeekableXZInputStream, SeekableXZInputStream, getCheckTypes, getIndexMemoryUsage, getLargestBlockSize, getStreamCount, getBlockCount, getBlockPos, getBlockSize, getBlockCompPos, getBlockCompSize, getBlockCheckType, getBlockNumber, read, read, available, close, length, position, seek, seekToBlock




Creates a new seekable XZ decompressor without a memory usage limit.
param   in seekable input stream containing one or more
XZ Streams; the whole input stream is used
throws   XZFormatException
input is not in the XZ format
throws   CorruptedInputException
XZ data is corrupt or truncated
throws   UnsupportedOptionsException
XZ headers seem valid but they specify
options not supported by this implementation
throws   EOFException
less than 6 bytes of input was available
from in, or (unlikely) the size
of the underlying stream got smaller while
this was reading from it
throws   IOException may be thrown by in
public void SeekableXZInputStream (SeekableInputStream in)


Creates a new seekable XZ decomporessor with an optional
memory usage limit.
param   in seekable input stream containing one or more
XZ Streams; the whole input stream is used
param   memoryLimit memory usage limit in kibibytes (KiB)
or -1 to impose no
memory usage limit
throws   XZFormatException
input is not in the XZ format
throws   CorruptedInputException
XZ data is corrupt or truncated
throws   UnsupportedOptionsException
XZ headers seem valid but they specify
options not supported by this implementation
throws   MemoryLimitException
decoded XZ Indexes would need more memory
than allowed by the memory usage limit
throws   EOFException
less than 6 bytes of input was available
from in, or (unlikely) the size
of the underlying stream got smaller while
this was reading from it
throws   IOException may be thrown by in
public void SeekableXZInputStream (SeekableInputStream in, int memoryLimit)


Gets the types of integrity checks used in the .xz file.
Multiple checks are possible only if there are multiple
concatenated XZ Streams.


The returned value has a bit set for every check type that is present.
For example, if CRC64 and SHA-256 were used, the return value is
(1 << XZ.CHECK_CRC64)
| (1 << XZ.CHECK_SHA256)
.

public int getCheckTypes ()


Gets the amount of memory in kibibytes (KiB) used by
the data structures needed to locate the XZ Blocks.
This is usually useless information but since it is calculated
for memory usage limit anyway, it is nice to make it available to too.
public int getIndexMemoryUsage ()


Gets the uncompressed size of the largest XZ Block in bytes.
This can be useful if you want to check that the file doesn't
have huge XZ Blocks which could make seeking to arbitrary offsets
very slow. Note that huge Blocks don't automatically mean that
seeking would be slow, for example, seeking to the beginning of
any Block is always fast.
public long getLargestBlockSize ()


Gets the number of Streams in the .xz file.
since  1.3
public int getStreamCount ()


Gets the number of Blocks in the .xz file.
since  1.3
public int getBlockCount ()


Gets the uncompressed start position of the given Block.

@throws IndexOutOfBoundsException if
blockNumber < 0 or
blockNumber >= getBlockCount().

@since 1.3
public long getBlockPos (int blockNumber)


Gets the uncompressed size of the given Block.

@throws IndexOutOfBoundsException if
blockNumber < 0 or
blockNumber >= getBlockCount().

@since 1.3
public long getBlockSize (int blockNumber)


Gets the position where the given compressed Block starts in
the underlying .xz file.
This information is rarely useful to the users of this class.

@throws IndexOutOfBoundsException if
blockNumber < 0 or
blockNumber >= getBlockCount().

@since 1.3
public long getBlockCompPos (int blockNumber)


Gets the compressed size of the given Block.
This together with the uncompressed size can be used to calculate
the compression ratio of the specific Block.

@throws IndexOutOfBoundsException if
blockNumber < 0 or
blockNumber >= getBlockCount().

@since 1.3
public long getBlockCompSize (int blockNumber)


Gets integrity check type (Check ID) of the given Block.

@throws IndexOutOfBoundsException if
blockNumber < 0 or
blockNumber >= getBlockCount().

@see #getCheckTypes()

@since 1.3
public int getBlockCheckType (int blockNumber)


Gets the number of the Block that contains the byte at the given
uncompressed position.

@throws IndexOutOfBoundsException if
pos < 0 or
pos >= length().

@since 1.3
public int getBlockNumber (long pos)


Decompresses the next byte from this input stream.
return   the next decompressed byte, or -1
to indicate the end of the compressed stream
throws   CorruptedInputException
throws   UnsupportedOptionsException
throws   MemoryLimitException
throws   XZIOException if the stream has been closed
throws   IOException may be thrown by in
public int read ()


Decompresses into an array of bytes.


If len is zero, no bytes are read and 0
is returned. Otherwise this will try to decompress len
bytes of uncompressed data. Less than len bytes may
be read only in the following situations:



@param buf target buffer for uncompressed data
@param off start offset in buf
@param len maximum number of uncompressed bytes to read

@return number of bytes read, or -1 to indicate
the end of the compressed stream

@throws CorruptedInputException
@throws UnsupportedOptionsException
@throws MemoryLimitException

@throws XZIOException if the stream has been closed

@throws IOException may be thrown by in
public int read (byte[] buf, int off, int len)


Returns the number of uncompressed bytes that can be read
without blocking. The value is returned with an assumption
that the compressed input data will be valid. If the compressed
data is corrupt, CorruptedInputException may get
thrown before the number of bytes claimed to be available have
been read from this input stream.
return   the number of uncompressed bytes that can be read
without blocking
public int available ()


Closes the stream and calls in.close().
If the stream was already closed, this does nothing.
throws   IOException if thrown by in.close()
public void close ()


Gets the uncompressed size of this input stream. If there are multiple
XZ Streams, the total uncompressed size of all XZ Streams is returned.
public long length ()


Gets the current uncompressed position in this input stream.
throws   XZIOException if the stream has been closed
public long position ()


Seeks to the specified absolute uncompressed position in the stream.
This only stores the new position, so this function itself is always
very fast. The actual seek is done when read is called
to read at least one byte.


Seeking past the end of the stream is possible. In that case
read will return -1 to indicate
the end of the stream.

@param pos new uncompressed read position

@throws XZIOException
if pos is negative, or
if stream has been closed

public void seek (long pos)


Seeks to the beginning of the given XZ Block.

@throws XZIOException
if blockNumber < 0 or
blockNumber >= getBlockCount(),
or if stream has been closed

@since 1.3
public void seekToBlock (int blockNumber)