libav.git - [no description]

	Commit message (Collapse)	Author	Age
*	x86: avcodec: Consistently name all init files	Diego Biurrun	2012-08-16
\|
*	Don't include common.h from avutil.h	Martin Storsjö	2012-08-15
\| \| \| \|	Signed-off-by: Martin Storsjö <martin@martin.st>
*	x86: avcodec: Appropriately name files containing only init functions	Diego Biurrun	2012-08-15
\|
*	mpegvideo_mmx_template: drop some commented-out cruft	Diego Biurrun	2012-08-15
\|
*	x86: cabac: allow building with suncc	Mans Rullgard	2012-08-13
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This fixes two issues preventing suncc from building this code. The undocumented 'a' operand modifier, causing gcc to omit a $ in front of immediate operands (as required in addresses), is not supported by suncc. Luckily, the also undocumented 'c' modifer has the same effect and is supported. On some asm statements with a large number of operands, suncc for no obvious reason fails to correctly substitute some of the operands. Fortunately, some of the operands in these statements are plain numbers which can be inserted directly into the code block instead of passed as operands. With these changes, the code builds correctly with both gcc and suncc. Signed-off-by: Mans Rullgard <mans@mansr.com>
*	x86: mlpdsp: avoid taking address of void	Mans Rullgard	2012-08-13
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	This code contains a C array of addresses of labels defined in inline asm. To do this, the names must be declared as external in C. The declared type does not matter since only the address is used, and for some reason, the author of the code used the 'void' type despite taking the address of a void expression being invalid. Changing the type to char, a reasonable choice since the alignment of the code labels cannot be known or guaranteed, eliminates gcc warnings and allows building with suncc. Signed-off-by: Mans Rullgard <mans@mansr.com>
*	x86: Drop silly "_yasm" suffixes from filenames	Diego Biurrun	2012-08-12
\|
*	Move MASK_ABS macro to libavcodec/mathops.h	Mans Rullgard	2012-08-09
\| \| \| \| \| \| \| \| \| \| \| \|	This macro is only used in two places, both in libavcodec, so this is a more sensible place for it. Two small tweaks to the macro are made: - removing the trailing semicolon - dropping unnecessary 'volatile' from the x86 asm Signed-off-by: Mans Rullgard <mans@mansr.com>
*	x86: rename libavutil/x86_cpu.h to libavutil/x86/asm.h	Mans Rullgard	2012-08-09
\| \| \| \| \| \| \|	This puts x86-specific things in the x86/ subdirectory where they belong. Signed-off-by: Mans Rullgard <mans@mansr.com>
*	x86: pngdsp: Fix assembly for OS/2	Dave Yeo	2012-08-08
\| \| \| \| \| \| \|	The a.out object format does not allow aligning sections. On OS/2 LD aligns sections to 16 bytes. Signed-off-by: Diego Biurrun <diego@biurrun.de>
*	x86: use 32-bit source registers with movd instruction	Mans Rullgard	2012-08-07
\| \| \| \| \| \| \| \|	yasm tolerates mismatch between movd/movq and source register size, adjusting the instruction according to the register. nasm is more strict. Signed-off-by: Mans Rullgard <mans@mansr.com>
*	x86: add colons after labels	Mans Rullgard	2012-08-07
\| \| \| \| \| \|	nasm prints a warning if the colon is missing. Signed-off-by: Mans Rullgard <mans@mansr.com>
*	Replace all CODEC_ID_* with AV_CODEC_ID_*	Anton Khirnov	2012-08-07
\|
*	x86: h264_idct: Rename x264_add8x4_idct_sse2 --> h264_add8x4_idct_sse2	Diego Biurrun	2012-08-05
\|
*	fft: 3dnow: fix register name typo in DECL_IMDCT macro	Ronald S. Bultje	2012-08-04
\| \| \| \|	Signed-off-by: Diego Biurrun <diego@biurrun.de>
*	x86: dct32: port to cpuflags	Diego Biurrun	2012-08-03
\|
*	x86: build: replace mmx2 by mmxext	Diego Biurrun	2012-08-03
\| \| \| \| \| \| \|	Refactoring mmx2/mmxext YASM code with cpuflags will force renames. So switching to a consistent naming scheme beforehand is sensible. The name "mmxext" is more official and widespread and also the name of the CPU flag, as reported e.g. by the Linux kernel.
*	dsputil: make add_hfyu_left_prediction_sse4() support unaligned src.	Ronald S. Bultje	2012-08-03
\| \| \| \| \| \| \| \| \| \|	This makes add_hfyu_left_prediction_sse4() handle sources that are not 16-byte aligned in its own function rather than by proxying the call to add_hfyu_left_prediction_ssse3(). This fixes a crash on Win64, since the sse4 version clobberes xmm6, but the ssse3 version (which uses MMX regs) does not restore it, thus leading to XMM clobbering and RSP being off. Fixes bug 342.
*	x86: Use consistent 3dnowext function and macro name suffixes	Diego Biurrun	2012-08-03
\| \| \| \| \| \|	Currently there is a wild mix of 3dn2/3dnow2/3dnowext. Switching to "3dnowext", which is a more common name of the CPU flag, as reported e.g. by the Linux kernel, unifies this.
*	x86: proresdsp: improve SIGNEXTEND macro comments	Diego Biurrun	2012-08-02
\|
*	x86: h264dsp: K&R formatting cosmetics	Diego Biurrun	2012-08-02
\|
*	x86: fft: fix imdct_half() for AVX	Ronald S. Bultje	2012-08-02
\| \| \| \| \| \| \| \| \|	Some calculations were changed in b6a3849 to use mmsize, which was not correct for the AVX version, which uses INIT_YMM and therefore has mmsize == 32. Fixes Bug 341. Signed-off-by: Justin Ruggles <justin.ruggles@gmail.com>
*	x86: remove libmpeg2 mmx(ext) idct functions	Mans Rullgard	2012-08-02
\| \| \| \| \| \| \| \|	These functions are not faster than other mmx implementations on any hardware I have been able to test on, and they are horribly inaccurate. There is thus no reason to ever use them. Signed-off-by: Mans Rullgard <mans@mansr.com>
*	fft: port FFT/IMDCT 3dnow functions to yasm, and disable on x86-64.	Ronald S. Bultje	2012-07-31
\| \| \| \| \|	64-bit CPUs always have SSE available, thus there is no need to compile in the 3dnow functions. This results in smaller binaries.
*	x86/dsputilenc: bury inline asm under HAVE_INLINE_ASM.	Ronald S. Bultje	2012-07-31
\|
*	x86: h264dsp: Remove unused variable ff_pb_3_1	Diego Biurrun	2012-08-01
\|
*	x86: h264dsp: Adjust YASM #ifdefs	Diego Biurrun	2012-07-31
\| \| \| \|	This fixes compilation with YASM disabled.
*	h264: convert loop filter strength dsp function to yasm.	Ronald S. Bultje	2012-07-30
\| \| \| \| \| \| \|	This completes the conversion of h264dsp to yasm; note that h264 also uses some dsputil functions, most notably qpel. Performance-wise, the yasm-version is ~10 cycles faster (182->172) on x86-64, and ~8 cycles faster (201->193) on x86-32.
*	h264_idct_10bit: port x86 assembly to cpuflags.	Ronald S. Bultje	2012-07-28
\|
*	fft: rename "z" to "zc" to prevent name collision.	Ronald S. Bultje	2012-07-28
\| \| \| \| \| \|	Without this, cglobal will expand "z" to "zh" to access the high byte in a register's word, which causes a name collision with the ZH(x) macro further up in this file.
*	vp3: don't compile mmx IDCT functions on x86-64.	Ronald S. Bultje	2012-07-27
\| \| \| \| \|	64-bit CPUs always have SSE2, and a SSE2 version exists, thus the MMX version will never be used.
*	h264_loopfilter: port x86 simd to cpuflags.	Ronald S. Bultje	2012-07-27
\|
*	h264_chromamc_10bit: port x86 simd to cpuflags.	Ronald S. Bultje	2012-07-27
\|
*	vp3: port x86 SIMD to cpuflags.	Ronald S. Bultje	2012-07-27
\|
*	rv34: port x86 SIMD to cpuflags.	Ronald S. Bultje	2012-07-27
\|
*	vp56: only compile MMX SIMD on x86-32.	Ronald S. Bultje	2012-07-27
\| \| \| \| \|	All x86-64 CPUs have SSE2, so the MMX version will never be used. This leads to smaller binaries.
*	vp56: port x86 simd to cpuflags.	Ronald S. Bultje	2012-07-27
\|
*	proresdsp: port x86 assembly to cpuflags.	Ronald S. Bultje	2012-07-27
\|
*	mpegaudio: bury inline asm under HAVE_INLINE_ASM.	Ronald S. Bultje	2012-07-26
\|
*	x86inc: automatically insert vzeroupper for YMM functions.	Ronald S. Bultje	2012-07-26
\|
*	vp3: don't use calls to inline asm in yasm code.	Ronald S. Bultje	2012-07-25
\| \| \| \| \| \| \| \|	Mixing yasm and inline asm is a bad idea, since if either yasm or inline asm is not supported by your toolchain, all of the asm stops working. Thus, better to use either one or the other alone. Signed-off-by: Derek Buitenhuis <derek.buitenhuis@gmail.com>
*	x86/dsputil: put inline asm under HAVE_INLINE_ASM.	Ronald S. Bultje	2012-07-25
\| \| \| \| \| \| \|	This allows compiling with compilers that don't support gcc-style inline assembly. Signed-off-by: Derek Buitenhuis <derek.buitenhuis@gmail.com>
*	dsputil_mmx: fix incorrect assembly code	Yang Wang	2012-07-25
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	In ff_put_pixels_clamped_mmx(), there are two assembly code blocks. In the first block (in the unrolled loop), the instructions "movq 8%3, %%mm1 \n\t", and so forth, have problems. From above instruction, it is clear what the programmer wants: a load from p + 8. But this assembly code doesn’t guarantee that. It only works if the compiler puts p in a register to produce an instruction like this: "movq 8(%edi), %mm1". During compiler optimization, it is possible that the compiler will be able to constant propagate into p. Suppose p = &x[10000]. Then operand 3 can become 10000(%edi), where %edi holds &x. And the instruction becomes "movq 810000(%edx)". That is, it will stride by 810000 instead of 8. This will cause a segmentation fault. This error was fixed in the second block of the assembly code, but not in the unrolled loop. How to reproduce: This error is exposed when we build using Intel C++ Compiler, with IPO+PGO optimization enabled. Crashed when decoding an MJPEG video. Signed-off-by: Michael Niedermayer <michaelni@gmx.at> Signed-off-by: Derek Buitenhuis <derek.buitenhuis@gmail.com>
*	dsputil: x86: add SHUFFLE_MASK_W macro	Jason Garrett-Glaser	2012-07-22
\| \| \| \|	Simplifies pshufb masks that operate on words.
*	x86: dsputil: drop some unused CPU flag debug code	Diego Biurrun	2012-07-19
\|
*	vp3: move idct and loop filter pointers to new vp3dsp context	Mans Rullgard	2012-07-18
\| \| \| \| \| \| \| \|	This moves all VP3-specific function pointers from dsputil to a new vp3dsp context. There is no reason to ever use the VP3 IDCT where an MPEG2 IDCT is expected or vice versa. Signed-off-by: Mans Rullgard <mans@mansr.com>
*	build: add CONFIG_VP3DSP, reduce repetition in OBJS lists	Mans Rullgard	2012-07-18
\| \| \| \|	Signed-off-by: Mans Rullgard <mans@mansr.com>
*	x86: h264_intrapred: Don't add the 'd' suffix to the SPLATB_REG macro	Martin Storsjö	2012-07-06
\| \| \| \| \| \| \| \| \| \| \| \| \|	The SPLATB_REG macro already adds the 'd' suffix internally. This fixes building on Win64, which has been broken since 878e66902. This worked for unix, where r2 happened to be rdx in this case, which with the first suffix rdxd was mapped to eax, and eaxd is defined back to eax. On win64 however, r2 happened to be R8 in this case, and R8d mapps to R8D just fine, but there's no mapping for R8Dd to anything. Signed-off-by: Martin Storsjö <martin@martin.st>
*	x86: h264_intrapred: use newly introduced SPLAT* and PSHUFLW macros	Diego Biurrun	2012-07-05
\|
*	x86inc: add SPLATB_LOAD, SPLATB_REG, PSHUFLW macros	Loren Merritt	2012-07-05
\| \| \| \|	Signed-off-by: Diego Biurrun <diego@biurrun.de>