The decompressor supports bigger dictionaries, up to almost 2 GiB.
With HC4 the encoder would support dictionaries bigger than 768 MiB.
The 768 MiB limit comes from the current implementation of BT4 where
we would otherwise hit the limits of signed ints in array indexing.
If you really need bigger dictionary for decompression,
use {@link LZMA2InputStream} directly.
niceLen is 8.niceLen is 273.LZMA2Options(PRESET_DEFAULT).preset is not supported
The presets 0-3 are fast presets with medium compression.
The presets 4-6 are fairly slow presets with high compression.
The default preset (PRESET_DEFAULT) is 6.
The presets 7-9 are like the preset 6 but use bigger dictionaries
and have higher compressor and decompressor memory requirements.
Unless the uncompressed size of the file exceeds 8 MiB,
16 MiB, or 32 MiB, it is waste of memory to use the
presets 7, 8, or 9, respectively.
@throws UnsupportedOptionsException
preset is not supported
The dictionary (or history buffer) holds the most recently seen
uncompressed data. Bigger dictionary usually means better compression.
However, using a dictioanary bigger than the size of the uncompressed
data is waste of memory.
Any value in the range [DICT_SIZE_MIN, DICT_SIZE_MAX] is valid,
but sizes of 2^n and 2^n + 2^(n-1) bytes are somewhat
recommended.
@throws UnsupportedOptionsException
dictSize is not supported
The .xz format doesn't support a preset dictionary for now.
Do not set a preset dictionary unless you use raw LZMA2.
Preset dictionary can be useful when compressing many similar,
relatively small chunks of data independently from each other.
A preset dictionary should contain typical strings that occur in
the files being compressed. The most probable strings should be
near the end of the preset dictionary. The preset dictionary used
for compression is also needed for decompression.
The sum of lc and lp is limited to 4.
Trying to exceed it will throw an exception. This function lets
you change both at the same time.
@throws UnsupportedOptionsException
lc and lp
are invalid
All bytes that cannot be encoded as matches are encoded as literals.
That is, literals are simply 8-bit bytes that are encoded one at
a time.
The literal coding makes an assumption that the highest lc
bits of the previous uncompressed byte correlate with the next byte.
For example, in typical English text, an upper-case letter is often
followed by a lower-case letter, and a lower-case letter is usually
followed by another lower-case letter. In the US-ASCII character set,
the highest three bits are 010 for upper-case letters and 011 for
lower-case letters. When lc is at least 3, the literal
coding can take advantage of this property in the uncompressed data.
The default value (3) is usually good. If you want maximum compression,
try setLc(4). Sometimes it helps a little, and sometimes it
makes compression worse. If it makes it worse, test for example
setLc(2) too.
@throws UnsupportedOptionsException
lc is invalid, or the sum
of lc and lp
exceed LC_LP_MAX
This affets what kind of alignment in the uncompressed data is
assumed when encoding literals. See {@link #setPb(int) setPb} for
more information about alignment.
@throws UnsupportedOptionsException
lp is invalid, or the sum
of lc and lp
exceed LC_LP_MAX
This affects what kind of alignment in the uncompressed data is
assumed in general. The default (2) means four-byte alignment
(2^pb = 2^2 = 4), which is often a good choice when
there's no better guess.
When the alignment is known, setting the number of position bits
accordingly may reduce the file size a little. For example with text
files having one-byte alignment (US-ASCII, ISO-8859-*, UTF-8), using
setPb(0) can improve compression slightly. For UTF-16
text, setPb(1) is a good choice. If the alignment is
an odd number like 3 bytes, setPb(0) might be the best
choice.
Even though the assumed alignment can be adjusted with
setPb and setLp, LZMA2 still slightly favors
16-byte alignment. It might be worth taking into account when designing
file formats that are likely to be often compressed with LZMA2.
@throws UnsupportedOptionsException
pb is invalid
This specifies the method to analyze the data produced by
a match finder. The default is MODE_FAST for presets
0-3 and MODE_NORMAL for presets 4-9.
Usually MODE_FAST is used with Hash Chain match finders
and MODE_NORMAL with Binary Tree match finders. This is
also what the presets do.
The special mode MODE_UNCOMPRESSED doesn't try to
compress the data at all (and doesn't use a match finder) and will
simply wrap it in uncompressed LZMA2 chunks.
@throws UnsupportedOptionsException
mode is not supported
niceLen bytes is found,
niceLen is invalid
Match finder has a major effect on compression speed, memory usage,
and compression ratio. Usually Hash Chain match finders are faster
than Binary Tree match finders. The default depends on the preset:
0-3 use MF_HC4 and 4-9 use MF_BT4.
@throws UnsupportedOptionsException
mf is not supported
The default is a special value of 0 which indicates that
the depth limit should be automatically calculated by the selected
match finder from the nice length of matches.
Reasonable depth limit for Hash Chain match finders is 4-100 and
16-1000 for Binary Tree match finders. Using very high values can
make the compressor extremely slow with some files. Avoid settings
higher than 1000 unless you are prepared to interrupt the compression
in case it is taking far too long.
@throws UnsupportedOptionsException
depthLimit is invalid
The returned value may bigger than the value returned by a direct call
to {@link LZMA2InputStream#getMemoryUsage(int)} if the dictionary size
is not 2^n or 2^n + 2^(n-1) bytes. This is because the .xz
headers store the dictionary size in such a format and other values
are rounded up to the next such value. Such rounding is harmess except
it might waste some memory if an unsual dictionary size is used.
If you use raw LZMA2 streams and unusual dictioanary size, call
{@link LZMA2InputStream#getMemoryUsage} directly to get raw decoder
memory requirements.
While this allows setting the LZMA2 compression options in detail,
often you only need
LZMA2Options()orLZMA2Options(int).