Commits · 8b8492452d53293b2ac8c842877fadf7925fc950 · Linshizhi / ffmpeg.wasm-core

30 Jul, 2017 1 commit

mdct15: add inverse transform postrotation SIMD · 70eb77b3

Rostislav Pehlivanov authored 7 years ago

2.5ms frames:
Before   (c):  2638 decicycles in postrotate, 2097040 runs,    112 skips
After (sse3):  1467 decicycles in postrotate, 2097083 runs,     69 skips
After (avx2):  1244 decicycles in postrotate, 2097085 runs,     67 skips

5ms frames:
Before   (c):  4987 decicycles in postrotate, 1048371 runs,    205 skips
After (sse3):  2644 decicycles in postrotate, 1048509 runs,     67 skips
After (avx2):  2031 decicycles in postrotate, 1048523 runs,     53 skips

10ms frames:
Before   (c):  9153 decicycles in postrotate,  523575 runs,    713 skips
After (sse3):  5110 decicycles in postrotate,  523726 runs,    562 skips
After (avx2):  3738 decicycles in postrotate,  524223 runs,     65 skips

20ms frames:
Before   (c): 17857 decicycles in postrotate,  261866 runs,    278 skips
After (sse3): 10041 decicycles in postrotate,  261746 runs,    398 skips
After (avx2):  7050 decicycles in postrotate,  262116 runs,     28 skips

Improves total decoding performance for real world content by 9% with avx2.
Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>

70eb77b3

11 Jul, 2017 1 commit

mdct15: remove redundant scale argument to imdct_half · aef5f9ab

Rostislav Pehlivanov authored 7 years ago

The only use of that argument was for Opus downmixing which is very rare
and better done after the mdcts.
Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>

aef5f9ab

23 Jun, 2017 1 commit

mdct15: add assembly optimizations for the 15-point FFT · e1120b1c

Rostislav Pehlivanov authored 7 years ago

c:    1802 decicycles in fft15,16774635 runs,   2581 skips
avx:   865 decicycles in fft15,16776378 runs,    838 skips
Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>

e1120b1c

14 Feb, 2017 1 commit

imdct15: rename to mdct15 and add a forward transform · d2119f62

Rostislav Pehlivanov authored 8 years ago

Handles strides (needed for Opus transients), does pre-reindexing and folding
without needing a copy.
Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>

d2119f62

05 Jan, 2017 2 commits

imdct15: replace the FFT with a faster PFA FFT algorithm · 2d208aaa

Rostislav Pehlivanov authored 8 years ago

This commit replaces the current inefficient non-power-of-two FFT with a
much faster FFT based on the Prime Factor Algorithm.
Although it is already much faster than the old algorithm without SIMD,
the new algorithm makes use of the already very throughouly SIMD'd power
of two FFT, which improves performance even more across all platforms
which we have SIMD support for.

Most of the work was done by Peter Barfuss, who passed the code to me to
implement into the iMDCT and the current codebase. The code for a
5-point and 15-point FFT was derived from the previous implementation,
although it was optimized and simplified, which will make its future
SIMD easier. The 15-point FFT is currently using 6% of the current
overall decoder overhead.

The FFT can now easily be used as a forward transform by simply not
multiplying the 5-point FFT's imaginary component by -1 (which comes
from the fact that changing the complex exponential's angle by -1 also
changes the output by that) and by multiplying the "theta" angle of the
main exptab by -1. Hence the deliberately left multiplication by -1 at
the end.

FATE passes, and performance reports on other platforms/CPUs are
welcome.

Performance comparisons:

iMDCT, PFA:
101127 decicycles in speed,   32765 runs,      3 skips
iMDCT, Old:
211022 decicycles in speed,   32768 runs,      0 skips

Standalone FFT, 300000 transforms of size 960:
    PFA        Old FFT     kiss_fft    libfftw3f
    3.659695s, 15.726912s, 13.300789s, 1.182222s

Being only 3x slower than libfftw3f is a big achievement by itself.

There appears to be something capping the performance in the iMDCT side
of things, possibly during the pre-stage reindexing. However, it is
certainly fast enough for now.
Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>

2d208aaa

imdct15: remove the AArch64 assembly · 4fdacf4c

Rostislav Pehlivanov authored 8 years ago

Prep work for the next commit, which will add a new FFT algorithm
which makes the iMDCT over 3x faster than it is currently (standalone,
the FFT is with some framesizes over 10x faster).

The new FFT algorithm uses the already thouroughly SIMD'd power of two
FFT which already has SIMD for AArch64, so users of that platform will
still see an improvement.

The previous FFT+SIMD was barely 2.5x faster than the C versions on these
platforms.
Signed-off-by: Rostislav Pehlivanov <atomnuker@gmail.com>

4fdacf4c

02 Feb, 2015 1 commit
- opus: Factor out imdct15 into a standalone component · 3d5d4623
  Diego Biurrun authored 10 years ago
```
It will be reused by the AAC decoder.
```
  3d5d4623
15 May, 2014 1 commit
- aarch64: opus NEON iMDCT and FFT · d3f5b947
  Janne Grunau authored 10 years ago
```
Opus celt decoding 11% faster and the iMDCT over 2.5 times faster on
Apple's A7.
```
  d3f5b947
22 Mar, 2014 1 commit
- dsputil: Refactor duplicated CALL_2X_PIXELS / PIXELS16 macros · 322a1dda
  Diego Biurrun authored 11 years ago
  
  322a1dda
13 Mar, 2014 1 commit
- rnd_avg.h: K&R formatting cosmetics · acd2b8e4
  Diego Biurrun authored 11 years ago
  
  acd2b8e4
19 Apr, 2013 1 commit
- hpeldsp: Add half-pel functions (currently copies of dsputil) · 68d8238c
  Ronald S. Bultje authored 11 years ago
```
Signed-off-by: Martin Storsjö <martin@martin.st>
```
  68d8238c
13 Mar, 2013 1 commit
- hpeldsp: add half-pel functions (currently copies of dsputil). · 9628e5a4
  Ronald S. Bultje authored 11 years ago
  
  9628e5a4
09 Feb, 2013 1 commit

rnd_avg: fix author attribution · 5fd6d85d

Michael Niedermayer authored 12 years ago

Reference:
commit 41fda91d
Author: BERO <bero@geocities.co.jp>
Date:   Wed May 14 17:46:55 2003 +0000

    aligned dsputil (for sh4) patch by (BERO <bero at geocities dot co dot jp>)

    Originally committed as revision 1880 to svn://svn.ffmpeg.org/ffmpeg/trunk

commit 8dbe5856
Author: Oskar Arvidsson <oskar@irock.se>
Date:   Tue Mar 29 17:48:59 2011 +0200

    Adds 8-, 9- and 10-bit versions of some of the functions used by the h264 decoder.

    This patch lets e.g. dsputil_init chose dsp functions with respect to
    the bit depth to decode. The naming scheme of bit depth dependent
    functions is <base name>_<bit depth>[_<prefix>] (i.e. the old
    clear_blocks_c is now named clear_blocks_8_c).

    Note: Some of the functions for high bit depth is not dependent on the
    bit depth, but only on the pixel size. This leaves some room for
    optimizing binary size.

    Preparatory patch for high bit depth h264 decoding support.
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

5fd6d85d

08 Feb, 2013 1 commit
- dsputil: Move rnd_avg inline functions to a separate header · bf6b3ec9
  Diego Biurrun authored 12 years ago
  
  bf6b3ec9
19 Mar, 2011 1 commit
- Replace FFmpeg with Libav in licence headers · 2912e87a
  Mans Rullgard authored 13 years ago
```
Signed-off-by: Mans Rullgard <mans@mansr.com>
```
  2912e87a
10 Jul, 2010 1 commit
- Add av_ prefix to bswap macros · 8fc0162a
  Måns Rullgård authored 14 years ago
```
Originally committed as revision 24170 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  8fc0162a
20 Apr, 2010 1 commit

Remove explicit filename from Doxygen @file commands. · ba87f080

Diego Biurrun authored 14 years ago

Passing an explicit filename to this command is only necessary if the
documentation in the @file block refers to a file different from the
one the block resides in.

Originally committed as revision 22921 to svn://svn.ffmpeg.org/ffmpeg/trunk

ba87f080

09 Mar, 2010 1 commit

Replace many includes of libavutil/common.h with what is actually needed · 2ed6f399

Måns Rullgård authored 14 years ago

This reduces the number of false dependencies on header files and
speeds up compilation.

Originally committed as revision 22407 to svn://svn.ffmpeg.org/ffmpeg/trunk

2ed6f399

01 Feb, 2009 1 commit

Use full internal pathname in doxygen @file directives. · bad5537e

Diego Biurrun authored 16 years ago

Otherwise doxygen complains about ambiguous filenames when files exist
under the same name in different subdirectories.

Originally committed as revision 16912 to svn://svn.ffmpeg.org/ffmpeg/trunk

bad5537e

21 Oct, 2008 1 commit
- split bswap.h into per-arch files · 3a90480a
  Måns Rullgård authored 16 years ago
```
Originally committed as revision 15663 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  3a90480a
31 Aug, 2008 1 commit

Globally rename the header inclusion guard names. · 98790382

Stefano Sabatini authored 16 years ago

Consistently apply this rule: the guard name is obtained from the
filename by stripping the leading "lib", converting '/' and '.'  to
'_' and uppercasing the resulting name. Guard names in the root
directory have to be prefixed by "FFMPEG_".

Originally committed as revision 15120 to svn://svn.ffmpeg.org/ffmpeg/trunk

98790382

09 May, 2008 1 commit
- Use full path for #includes from another directory. · 245976da
  Diego Biurrun authored 16 years ago
```
Originally committed as revision 13098 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  245976da
17 Oct, 2007 1 commit
- Add FFMPEG_ prefix to all multiple inclusion guards. · 5b21bdab
  Diego Biurrun authored 17 years ago
```
Originally committed as revision 10765 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  5b21bdab
20 May, 2007 2 commits
- add a ff_ prefix to some mpegaudio funcs · ca6e50af
  Aurelien Jacobs authored 17 years ago
```
Originally committed as revision 9081 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  ca6e50af
- loosen dependencies over mpegaudiodec · 4bd8e17c
  Aurelien Jacobs authored 17 years ago
```
Originally committed as revision 9080 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  4bd8e17c
07 Oct, 2006 1 commit
- Change license headers to say 'FFmpeg' instead of 'this program/this library' · b78e7197
  Diego Biurrun authored 18 years ago
```
and fix GPL/LGPL version mismatches.

Originally committed as revision 6577 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  b78e7197
19 Sep, 2006 1 commit
- New single instruction math operation header · 99aed7c8
  Luca Barbato authored 18 years ago
```
Originally committed as revision 6291 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  99aed7c8
12 Jan, 2006 1 commit
- Update licensing information: The FSF changed postal address. · 5509bffa
  Diego Biurrun authored 19 years ago
```
Originally committed as revision 4842 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  5509bffa
26 May, 2005 2 commits

cleanup · 3f3f8b2b

Michael Niedermayer authored 19 years ago

Originally committed as revision 4312 to svn://svn.ffmpeg.org/ffmpeg/trunk

3f3f8b2b

Better ARM support for mplayer/ffmpeg, ported from atty fork · 6ad1fa5a

Bernhard Rosenkränzer authored 19 years ago

while playing with some new hardware, I found it's running a forked mplayer
 -- and it looks like they're following the GPL.

 The maintainer's page is here: http://atty.jp/?Zaurus/mplayer
 Unfortunately it's mostly in Japanese, so it's hard to figure out any
  details.

  Their code looks quite interesting (at least to those of us w/ ARM CPUs).

  The patches I've attached are the patches from atty.jp with a couple of
  modifications by myself:
  - ported to current CVS
  - reverted their change of removing SNOW support from ffmpeg
  - cleaned up their bswap mess
  - removed DOS-style linebreaks from various files

patch by (Bernhard Rosenkraenzer: bero, arklinux org)

Originally committed as revision 4311 to svn://svn.ffmpeg.org/ffmpeg/trunk

6ad1fa5a

03 Mar, 2003 1 commit
- MpegEncContext.(i)dct_* -> DspContext.(i)dct_* · b0368839
  Michael Niedermayer authored 21 years ago
```
bitexact cleanup

Originally committed as revision 1617 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  b0368839
11 Feb, 2003 1 commit
- * UINTX -> uintx_t INTX -> intx_t · 0c1a9eda
  Zdenek Kabelac authored 22 years ago
```
Originally committed as revision 1578 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  0c1a9eda
20 Nov, 2002 1 commit

* cut&paste fix · bb285683

Zdenek Kabelac authored 22 years ago

Originally committed as revision 1249 to svn://svn.ffmpeg.org/ffmpeg/trunk

bb285683

19 Nov, 2002 2 commits

* oops fixed bad initialization of ff vals. · 59402627

Zdenek Kabelac authored 22 years ago

  - put FF_LIBMPEG2_IDCT_PERM into CVS - so it will work for now

Originally committed as revision 1227 to svn://svn.ffmpeg.org/ffmpeg/trunk

59402627

* compilation fix (ARM users please check) · 83f238cb
Zdenek Kabelac authored 22 years ago
```
Originally committed as revision 1225 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
83f238cb

25 Oct, 2002 1 commit
- idct_permutation_type variable, so the permutation type can quickly be identified · 50eb9cbc
  Michael Niedermayer authored 22 years ago
```
Originally committed as revision 1071 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  50eb9cbc
06 Oct, 2002 1 commit
- trying to fix the non-x86 IDCTs (untested) · 676e200c
  Michael Niedermayer authored 22 years ago
```
Originally committed as revision 1006 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  676e200c
25 May, 2002 1 commit
- license/copyright change · ff4ec49e
  Fabrice Bellard authored 22 years ago
```
Originally committed as revision 599 to svn://svn.ffmpeg.org/ffmpeg/trunk
```
  ff4ec49e
13 Aug, 2001 1 commit

arm specific code · 92651f67

Fabrice Bellard authored 23 years ago

Originally committed as revision 79 to svn://svn.ffmpeg.org/ffmpeg/trunk

92651f67