libav.git - [no description]

	Commit message (Collapse)	Author	Age
*	motion_est: convert stride to ptrdiff_t	Vittorio Giovara	2014-11-24
\| \| \| \| \|	CC: libav-stable@libav.org Bug-Id: CID 700556 / CID 700557 / CID 700558
*	idctdsp: Add global function pointers for {add\|put}_pixels_clamped functions	Diego Biurrun	2014-09-02
\| \| \| \| \| \|	These function pointers already existed in the ARM code. Adding them globally allows calls to the function pointers to access arch-optimized versions of the functions transparently.
*	build: Add explanatory comments to (optimization) blocks in the Makefiles	Diego Biurrun	2014-08-15
\|
*	mpegvideo: cosmetics: Lowercase ugly uppercase MPV_ function name prefixes	Diego Biurrun	2014-08-15
\|
*	vc-1: Add platform-specific start code search routine to VC1DSPContext.	Ben Avison	2014-08-04
\| \| \| \| \| \| \|	Initialise VC1DSPContext for parser as well as for decoder. Note, the VC-1 code doesn't actually use the function pointer yet. Signed-off-by: Luca Barbato <lu_zero@gentoo.org>
*	h264: Move start code search functions into separate source files.	Ben Avison	2014-08-04
\| \| \| \| \| \|	This permits re-use with parsers for codecs which use similar start codes. Signed-off-by: Luca Barbato <lu_zero@gentoo.org>
*	qpeldsp: Mark source pointer in qpel_mc_func function pointer const	Diego Biurrun	2014-07-25
\|
*	arm: Macroize the test for 'setend' CPU instruction support	Ben Avison	2014-07-21
\| \| \| \|	Signed-off-by: Diego Biurrun <diego@biurrun.de>
*	dct-test: Move arch-specific bits into arch-specific subdirectories	Diego Biurrun	2014-07-21
\|
*	idct: Move arm-specific declarations to a header in the arm directory	Diego Biurrun	2014-07-20
\|
*	idctdsp: prettyprinting cosmetics	Diego Biurrun	2014-07-18
\|
*	idct: Convert IDCT permutation #defines to an enum	Diego Biurrun	2014-07-18
\| \| \| \|	Also rename the enum values to be consistent with other DCT permutations.
*	arm: cosmetics: Consistently use lowercase for shift operators	Martin Storsjö	2014-07-18
\| \| \| \|	Signed-off-by: Martin Storsjö <martin@martin.st>
*	arm: cosmetics: Fix a misaligned asm operand	Martin Storsjö	2014-07-18
\| \| \| \|	Signed-off-by: Martin Storsjö <martin@martin.st>
*	armv6: Accelerate ff_fft_calc for general case (nbits != 4)	Ben Avison	2014-07-18
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The previous implementation targeted DTS Coherent Acoustics, which only requires nbits == 4 (fft16()). This case was (and still is) linked directly rather than being indirected through ff_fft_calc_vfp(), but now the full range from radix-4 up to radix-65536 is available. This benefits other codecs such as AAC and AC3. The implementaion is based upon the C version, with each routine larger than radix-16 calling a hierarchy of smaller FFT functions, then performing a post-processing pass. This pass benefits a lot from loop unrolling to counter the long pipelines in the VFP. A relaxed calling standard also reduces the overhead of the call hierarchy, and avoiding the excessive inlining performed by GCC probably helps with I-cache utilisation too. I benchmarked the result by measuring the number of gperftools samples that hit anywhere in the AAC decoder (starting from aac_decode_frame()) or specifically in the FFT routines (fft4() to fft512() and pass()) for the same sample AAC stream: Before After Mean StdDev Mean StdDev Confidence Change Audio decode 2245.5 53.1 1599.6 43.8 100.0% +40.4% FFT routines 940.6 22.0 348.1 20.8 100.0% +170.2% Signed-off-by: Martin Storsjö <martin@martin.st>
*	armv6: Accelerate ff_imdct_half for general case (mdct_bits != 6)	Ben Avison	2014-07-18
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The previous implementation targeted DTS Coherent Acoustics, which only requires mdct_bits == 6. This relatively small size lent itself to unrolling the loops a small number of times, and encoding offsets calculated at assembly time within the load/store instructions of each iteration. In the more general case (codecs such as AAC and AC3) much larger arrays are used - mdct_bits == [8, 9, 11]. The old method does not scale for these cases, so more integer registers are used with non-unrolled versions of the loops (and with some stack spillage). The postrotation filter loop is still unrolled by a factor of 2 to permit the double-buffering of some VFP registers to facilitate overlap of neighbouring iterations. I benchmarked the result by measuring the number of gperftools samples that hit anywhere in the AAC decoder (starting from aac_decode_frame()) or specifically in ff_imdct_half_c / ff_imdct_half_vfp, for the same example AAC stream: Before After Mean StdDev Mean StdDev Confidence Change aac_decode_frame 2368.1 35.8 2117.2 35.3 100.0% +11.8% ff_imdct_half_* 457.5 22.4 251.2 16.2 100.0% +82.1% Signed-off-by: Martin Storsjö <martin@martin.st>
*	dsputil: Split motion estimation compare bits off into their own context	Diego Biurrun	2014-07-17
\|
*	arm: dsputil: Coalesce all init files	Diego Biurrun	2014-07-16
\|
*	dsputil: Drop unused bit_depth parameter from all init functions	Diego Biurrun	2014-07-11
\|
*	dsputil: Split off pixel block routines into their own context	Diego Biurrun	2014-07-09
\|
*	arm: Avoid using the 'setend' instruction on ARMv7 and newer	Martin Storsjö	2014-07-08
\| \| \| \| \| \| \| \| \| \|	This instruction is deprecated on ARMv8, and it is serializing on some ARMv7 cores as well [1]. [1] http://article.gmane.org/gmane.linux.ports.arm.kernel/339293 CC: libav-stable@libav.org Signed-off-by: Martin Storsjö <martin@martin.st>
*	dsputil: Move pix_sum, pix_norm1, shrink function pointers to mpegvideoenc	Diego Biurrun	2014-07-06
\|
*	dsputil: Split off IDCT bits into their own context	Diego Biurrun	2014-06-30
\|
*	h264: avoid using uninitialized memory in NEON chroma mc	Janne Grunau	2014-06-23
\| \| \| \| \|	Adapt commit 982b596ea6640bfe218a31f6c3fc542d9fe61c31 for the arm and aarch64 NEON asm. 5-10% faster on Cortex-A9.
*	dsputil: Split audio operations off into a separate context	Diego Biurrun	2014-06-22
\|
*	dsputil: Split clear_block/fill_block off into a separate context	Diego Biurrun	2014-06-18
\|
*	arm: check if AS supports .dn	Janne Grunau	2014-06-03
\| \| \| \| \| \| \| \| \| \| \| \|	Move the GNU as check before the arch specific asm checks since the .dn check requires gas compatible assembler. Disable the VC-1 motion compensation NEON asm which is the only part using that directive. The integrated assembler in the upcoming clang 3.5 does not support .dn/.qn without plans to change that. Too much effort to implement it while it is rarely used. http://llvm.org/bugs/show_bug.cgi?id=18199.
*	dsputil: Move APE-specific bits into apedsp	Diego Biurrun	2014-05-29
\|
*	mpegvideo: move the MpegEncContext fields used from arm asm to the beginning	Anton Khirnov	2014-04-29
\| \| \| \| \|	This should reduce the frequency with which the offsets need to be updated.
*	lavu: add CHK_OFFS as AV_CHECK_OFFSET to check struct member offsets	Janne Grunau	2014-04-24
\|
*	Remove a number of unnecessary dsputil.h #includes	Diego Biurrun	2014-04-04
\|
*	arm: asm decode_block_coeffs_internal is vp8 specific	Janne Grunau	2014-04-04
\| \| \| \| \|	Unbreaks compilation on arm due to conflicting types for 'ff_decode_block_coeffs_armv6'.
*	On2 VP7 decoder	Peter Ross	2014-04-04
\| \| \| \| \| \| \| \| \|	Further performance improvements and security fixes by Vittorio Giovara, Luca Barbato and Diego Biurrun. Signed-off-by: Vittorio Giovara <vittorio.giovara@gmail.com> Signed-off-by: Luca Barbato <lu_zero@gentoo.org> Signed-off-by: Diego Biurrun <diego@biurrun.de>
*	arm: build: Maintain decoder objects separate from infrastructure objects	Diego Biurrun	2014-03-27
\|
*	truehd: add hand-scheduled ARM asm version of ff_mlp_pack_output.	Ben Avison	2014-03-26
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Profiling results for overall decode and the output_data function in particular are as follows: Before After Mean StdDev Mean StdDev Confidence Change 6:2 total 339.6 15.1 329.3 16.0 95.8% +3.1% (insignificant) 6:2 function 24.6 6.0 9.9 3.1 100.0% +148.5% 8:2 total 324.5 15.5 323.6 14.3 15.2% +0.3% (insignificant) 8:2 function 20.4 3.9 9.9 3.4 100.0% +104.7% 6:6 total 572.8 20.6 539.9 24.2 100.0% +6.1% 6:6 function 54.5 5.6 16.0 3.8 100.0% +240.9% 8:8 total 741.5 21.2 702.5 18.5 100.0% +5.6% 8:8 function 63.9 7.6 18.4 4.8 100.0% +247.3% The assembly version has also been tested with a fuzz tester to ensure that any combinations of inputs not exercised by my available test streams still generate mathematically identical results to the C version. Signed-off-by: Martin Storsjö <martin@martin.st>
*	truehd: add hand-scheduled ARM asm version of ff_mlp_rematrix_channel.	Ben Avison	2014-03-26
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Profiling results for overall audio decode and the rematrix_channels function in particular are as follows: Before After Mean StdDev Mean StdDev Confidence Change 6:2 total 370.8 17.0 348.8 20.1 99.9% +6.3% 6:2 function 46.4 8.4 45.8 6.6 18.0% +1.2% (insignificant) 8:2 total 343.2 19.0 339.1 15.4 54.7% +1.2% (insignificant) 8:2 function 38.9 3.9 40.2 6.9 52.4% -3.2% (insignificant) 6:6 total 658.4 15.7 604.6 20.8 100.0% +8.9% 6:6 function 109.0 8.7 59.5 5.4 100.0% +83.3% 8:8 total 896.2 24.5 766.4 17.6 100.0% +16.9% 8:8 function 223.4 12.8 93.8 5.0 100.0% +138.3% The assembly version has also been tested with a fuzz tester to ensure that any combinations of inputs not exercised by my available test streams still generate mathematically identical results to the C version. Signed-off-by: Martin Storsjö <martin@martin.st>
*	truehd: add hand-scheduled ARM asm version of mlp_filter_channel.	Ben Avison	2014-03-26
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Profiling results for overall audio decode and the mlp_filter_channel(_arm) function in particular are as follows: Before After Mean StdDev Mean StdDev Confidence Change 6:2 total 380.4 22.0 370.8 17.0 87.4% +2.6% (insignificant) 6:2 function 60.7 7.2 36.6 8.1 100.0% +65.8% 8:2 total 357.0 17.5 343.2 19.0 97.8% +4.0% (insignificant) 8:2 function 60.3 8.8 37.3 3.8 100.0% +61.8% 6:6 total 717.2 23.2 658.4 15.7 100.0% +8.9% 6:6 function 140.4 12.9 81.5 9.2 100.0% +72.4% 8:8 total 981.9 16.2 896.2 24.5 100.0% +9.6% 8:8 function 193.4 15.0 103.3 11.5 100.0% +87.2% Experiments with adding preload instructions to this function yielded no useful benefit, so these have not been included. The assembly version has also been tested with a fuzz tester to ensure that any combinations of inputs not exercised by my available test streams still generate mathematically identical results to the C version. Signed-off-by: Martin Storsjö <martin@martin.st>
*	dsputil: Refactor duplicated CALL_2X_PIXELS / PIXELS16 macros	Diego Biurrun	2014-03-22
\|
*	dsputil: Use correct type in me_cmp_func function pointer	Diego Biurrun	2014-03-20
\|
*	build: Group general components separate from de/encoders in arch Makefiles	Diego Biurrun	2014-03-20
\| \| \| \|	This is in line with how the top-level libavcodec Makefile is structured.
*	dsputil: Propagate bit depth information to all (sub)init functions	Diego Biurrun	2014-03-20
\| \| \| \|	This avoids recalculating the value over and over again.
*	arm: dsputil: K&R formatting cosmetics	Diego Biurrun	2014-03-20
\|
*	arm: dsputil: Drop restrict keyword from add_pixels_clamped_armv6 prototype	Diego Biurrun	2014-03-14
\| \| \| \| \| \| \| \|	The function is assigned to a function pointer that does not have the restrict keyword for that parameter. This fixes compilation for MSVC builds that don't recognize "restrict", broken since ed9625eb62.
*	Update dsputil- and SIMD-related comments to match reality more closely	Diego Biurrun	2014-03-13
\|
*	arm: dsputil: Add a bunch of missing #includes	Diego Biurrun	2014-03-13
\|
*	dsputil: Remove prototypes for nonexisting optimization functions	Diego Biurrun	2014-03-13
\|
*	armv6: vp8: use explicit labels in motion compensation asm	Janne Grunau	2014-03-12
\| \| \| \| \|	The integrated arm assembler in clang-503.0.38 (Xcode-5.1) fails to assemble a branch to 'label + offset' in thumb mode.
*	arm: get_cabac inline asm	Janne Grunau	2014-03-09
\| \| \| \| \| \| \| \| \| \|	Based on the aarch64 asm. CPU cycle counts on cortex-a9 compared to gcc 4.8.2: before: 475 decicycles in get_cabac_noinline, 67106035 runs, 2829 skips after: 393 decicycles in get_cabac_noinline, 67106474 runs, 2390 skips Overall speedup is above 2%. Code generated by clang 3.4 is slower on the same hardware and the relative change is a little larger.
*	arm: vp3: remove incorrect const in ff_vp3_idct_dc_add_neon declaration	Janne Grunau	2014-03-09
\| \| \| \| \|	Was missed in aeaf268e52fc11c1f64914a319e0edddf1346d6a when integrating clear_blocks into the idct.
*	arm: hpeldsp: fix put_pixels8_y2_{,no_rnd_}armv6	Janne Grunau	2014-03-08
\| \| \| \| \| \| \| \|	The overread avoidance fix in cbddee1cca0ebd01e8c5aa694d31228eb4de4b41 broke the computation for the last row since it prevented the safe reading from the height+1-th row. CC: libav-stable@libav.org