Commits · cf491a925e221122f81873bd041c5c136027e385 · Linshizhi / ffmpeg.wasm-core

09 Nov, 2015 10 commits

swresample/resample: speed up Blackman Nuttall filter · cf491a92

Ganesh Ajjanagadde authored Nov 09, 2015

This may be a slightly surprising optimization, but is actually based on
an understanding of how math libraries compute trigonometric functions.
Explanation is given here so that future development uses libm more effectively
across the codebase.

All libm's essentially compute transcendental functions via some kind of
polynomial approximation, be it Taylor-Maclaurin or Chebyshev.
Correction terms are added via polynomial correction factors when needed
to squeeze out the last bits of accuracy. Lookup tables are also
inserted strategically.

In the case of trigonometric functions, periodicity is exploited via
first doing a range reduction to an interval around zero, and then using
some polynomial approximation.

This range reduction is the most natural way of doing things - else one
would need polynomials for ranges in different periods which makes no
sense whatsoever.

To avoid the need for the range reduction, it is helpful to feed in
arguments as close to the origin as possible for the trigonometric
functions. In fact, this also makes sense from an accuracy point of view:
IEEE floating point has far more resolution for small numbers than big ones.

This patch does this for the Blackman-Nuttall filter, and yields a
non-negligible speedup.

Sample benchmark (x86-64, Haswell, GNU/Linux)
test: fate-swr-resample-dblp-2626-44100
old:
18893514 decicycles in build_filter (loop 1000),     256 runs,      0 skips
18599863 decicycles in build_filter (loop 1000),     512 runs,      0 skips
18445574 decicycles in build_filter (loop 1000),    1000 runs,     24 skips

new:
16290697 decicycles in build_filter (loop 1000),     256 runs,      0 skips
16267172 decicycles in build_filter (loop 1000),     512 runs,      0 skips
16251105 decicycles in build_filter (loop 1000),    1000 runs,     24 skips
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

cf491a92

swresample/resample: speed up upsampling by precomputing sines · b87ca4bf

Ganesh Ajjanagadde authored Nov 09, 2015

When upsampling, factor is set to 1 and sines need to be evaluated only
once for each phase, and the complexity should not depend on the number
of filter taps. This does the desired precomputation, yielding
significant speedups. Hard guarantees on the gain are not possible, but gains
themselves are obvious and are illustrated below.

Sample benchmark (x86-64, Haswell, GNU/Linux)
test: fate-swr-resample-dblp-2626-44100
old:
29161085 decicycles in build_filter (loop 1000),     256 runs,      0 skips
28821467 decicycles in build_filter (loop 1000),     512 runs,      0 skips
28668201 decicycles in build_filter (loop 1000),    1000 runs,     24 skips

new:
14351936 decicycles in build_filter (loop 1000),     256 runs,      0 skips
14306652 decicycles in build_filter (loop 1000),     512 runs,      0 skips
14299923 decicycles in build_filter (loop 1000),    1000 runs,     24 skips

Note that this does not statically allocate the sin lookup table. This
may be done for the default 1024 phases, yielding a 512*8 = 4kB array
which should be small enough.
This should yield a small improvement. Nevertheless, this is separate from
this patch, is more ambiguous due to the binary increase, and requires a
lut to be generated offline.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

b87ca4bf

doc/ffmpeg: Clarify that the sdp_file option requires an rtp output. · b02201ef
Simon Thelen authored Nov 02, 2015
```
Signed-off-by: Simon Thelen <ffmpeg-dev@c-14.de>
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
b02201ef

ffmpeg: Don't try and write sdp info if none of the outputs had an rtp format. · 70fb5ead

Simon Thelen authored Nov 02, 2015

Fixes a segfault when trying to write nonexistent rtp information.
Signed-off-by: Simon Thelen <ffmpeg-dev@c-14.de>
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>

70fb5ead

avformat/cache: Avoid int-overflow in cache compare function · 72f9a634

Bryan Huh authored Nov 09, 2015

cache protocol indexes its cache using AVTreeNodes which require a cmp
function for inserting and searching new cache-entries. This cmp
function expects a 32-bit int return value (negative, zero, or positive)
but the cache cmp function returns an int64_t which can overflow the
int, giving negative numbers for when it should be positive, vice versa.
This manifests itself only for very large files (e.g. 4GB+)
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>

72f9a634

avcodec/nvenc: update nvenc default parameters · ddbad158
Agatha Hu authored Nov 09, 2015
```
Signed-off-by: Timo Rothenpieler <timo@rothenpieler.org>
```
ddbad158
avutil/x86/intmath: Correct intrinsic headers for older compilers. · f9841745
Matt Oliver authored Nov 09, 2015
```
Signed-off-by: Matt Oliver <protogonoi@gmail.com>
```
f9841745
avformat/rsd: add XMA support · 0cfd4a99
Paul B Mahol authored Nov 09, 2015
```
Signed-off-by: Paul B Mahol <onemda@gmail.com>
```
0cfd4a99

swresample/resample: improve bessel function accuracy and speed · a5202bc9

Ganesh Ajjanagadde authored Nov 02, 2015

This improves accuracy for the bessel function at large arguments, and this in turn
should improve the quality of the Kaiser window. It also improves the
performance of the bessel function and hence build_filter by ~ 20%.
Details are given below.

Algorithm: taken from the Boost project, who have done a detailed
investigation of the accuracy of their method, as compared with e.g the
GNU Scientific Library (GSL):
http://www.boost.org/doc/libs/1_52_0/libs/math/doc/sf_and_dist/html/math_toolkit/special/bessel/mbessel.html.
Boost source code (also cited and licensed in the code):
https://searchcode.com/codesearch/view/14918379/.

Accuracy: sample values may be obtained as follows. i0 denotes the old bessel code,
i0_boost the approach here, and i0_real an arbitrary precision result (truncated) from Wolfram Alpha:
type "bessel i0(6.0)" to reproduce. These are evaluation points that occur for
the default kaiser_beta = 9.

Some illustrations:
bessel(8.0)
i0      (8.000000) = 427.564115721804739678191254
i0_boost(8.000000) = 427.564115721804796521610115
i0_real (8.000000) = 427.564115721804785177396791

bessel(6.0)
i0      (6.000000) = 67.234406976477956163762428
i0_boost(6.000000) = 67.234406976477970374617144
i0_real (6.000000) = 67.234406976477975326188025

Reason for accuracy: Main accuracy benefits come at larger bessel arguments, where the
Taylor-Maclaurin method is not that good: 23+ iterations
(at large arguments, since the series is about 0) can cause
significant floating point error accumulation.

Benchmarks: Obtained on x86-64, Haswell, GNU/Linux via a loop calling
build_filter 1000 times:
test: fate-swr-resample-dblp-44100-2626

new:
995894468 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1029719302 decicycles in build_filter(loop 1000),     512 runs,      0 skips
984101131 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

old:
1250020763 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1246353282 decicycles in build_filter(loop 1000),     512 runs,      0 skips
1220017565 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

A further ~ 5% may be squeezed by enabling -ftree-vectorize. However,
this is a separate issue from this patch.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

a5202bc9

swresample: allow double precision beta value for the Kaiser window · 1bed09a3

Ganesh Ajjanagadde authored Nov 07, 2015

Kaiser windows inherently don't require beta to be an integer. This was
an arbitrary restriction. Moreover, soxr does not require it, and in
fact often estimates beta to a non-integral value.

Thus, this patch allows greater flexibility for swresample clients.
Micro version is updated.
Reviewed-by: Derek Buitenhuis <derek.buitenhuis@gmail.com>
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

1bed09a3

08 Nov, 2015 19 commits
- softfloat: handle INT_MIN correctly in av_int2sf · 9ac61e73
  Andreas Cadhalpun authored Nov 08, 2015
```
Otherwise v=INT_MIN doesn't get normalized and thus triggers av_assert2
in other functions.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Andreas Cadhalpun <Andreas.Cadhalpun@googlemail.com>
```
  9ac61e73
- softfloat: assert when the argument of av_sqrt_sf is negative · f3866a14
  Andreas Cadhalpun authored Nov 08, 2015
```
The correct result can't be expressed in SoftFloat.
Currently it returns a random value from an out of bounds read.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Andreas Cadhalpun <Andreas.Cadhalpun@googlemail.com>
```
  f3866a14
- avfilter: add anoisesrc · 6a11c7f1
  Kyle Swanson authored Nov 08, 2015
```
Signed-off-by: Kyle Swanson <k@ylo.ph>
Signed-off-by: Paul B Mahol <onemda@gmail.com>
```
  6a11c7f1
- avutil/softfloat: Include negative numbers in cmp/gt tests · 955cdc43
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  955cdc43
- avutil/softfloat: Fix av_gt_sf() with large exponents try #2 · 05b05a7a
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  05b05a7a
- avutil/softfloat: Add test for av_gt_sf() · 791ea23e
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  791ea23e
- avutil/softfloat: Extend the av_cmp_sf() test to cover a wider range of exponents · ecfb0761
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  ecfb0761
- avutil/softfloat: Fix overflows in shifts in av_cmp_sf() and av_gt_sf() · cee3c9d2
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  cee3c9d2
- avutil/softfloat: Add test for av_cmp_sf() · df2a2117
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  df2a2117
- avutil/softfloat: Add tests for exponent underflows · 596dfe7d
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  596dfe7d
- avutil/softfloat: Fix exponent underflow in av_div_sf() · 046218b2
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  046218b2
- avutil/softfloat: Fix exponent underflow in av_mul_sf() · a1e3303f
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  a1e3303f
- avutil/softfloat: Fix typo in av_mul_sf() doxy · 4135a2bf
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  4135a2bf
- Revert "avutil/softfloat: Check for MIN_EXP in av_sqrt_sf()" · 4b6ad236
  Michael Niedermayer authored Nov 08, 2015
```
This case should not be possible if the input has a exponent within
the valid range

This reverts commit 0269fb11.
```
  4b6ad236
- avutil/softfloat: Check for MIN_EXP in av_sqrt_sf() · 0269fb11
  Michael Niedermayer authored Nov 08, 2015
```
Otherwise the exponent could eventually underflow
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  0269fb11
- avutil/softfloat: Correctly set the exponent for 0.0 in av_sqrt_sf() · 107db5ab
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  107db5ab
- avcodec/aacsbr: Use FLOAT_0 · dcf1cf5d
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  dcf1cf5d
- avutil/softfloat: FLOAT_0 should use MIN_EXP · a66b243d
  Michael Niedermayer authored Nov 08, 2015
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  a66b243d
- Add pixblockdsp checkasm tests · 3d20f8e7
  Timothy Gu authored Nov 01, 2015
  
  3d20f8e7
07 Nov, 2015 11 commits
- pixblockdsp: x86: Condense diff_pixels_* to a shared macro · 4b80b895
  Timothy Gu authored Nov 01, 2015
```
Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com>
Reviewed-by: James Almer <jamrial@gmail.com>
```
  4b80b895
- avcodec/takdec: Use memove, avoid undefined memcpy() use · 7cea3430
  Michael Niedermayer authored Nov 07, 2015
```
Fixes: e214333cbd94c91228e624ff39329ce6/asan_generic_4a5159_6412_96cda2530e80607210ab41ccae3d456d.tak

Found-by: Mateusz "j00ru" Jurczyk and Gynvael Coldwind
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  7cea3430
- mmaldec: correct package buffering accounting · a55fbfa4
  wm4 authored Nov 06, 2015
```
The assert in ffmmal_stop_decoder() could trigger sometimes. The
packets_buffered counter was indeed not correctly maintained, and
packets were not subtracted from it if they were still in the waiting
queue.

For some reason, this happened especially with VC-1.
```
  a55fbfa4
- mmaldec: add vc1 decoding support · b07cbf67
  wm4 authored Nov 06, 2015
  
  b07cbf67
- lavfi/af_asyncts: remove looping on request_frame(). · 785ac437
  Nicolas George authored Oct 24, 2015
  
  785ac437
- lavfi/af_amix: mostly fix scheduling. · a08fb398
  Nicolas George authored Oct 24, 2015
  
  a08fb398
- lavfi/vf_framepack: fix scheduling. · f53c4b6a
  Nicolas George authored Oct 22, 2015
  
  f53c4b6a
- lavfi/af_join: partially fix scheduling. · d0b82d79
  Nicolas George authored Oct 22, 2015
  
  d0b82d79
- lavfi/fifo: do not assume request_frame() returns a frame. · 67d3f529
  Nicolas George authored Oct 22, 2015
  
  67d3f529
- lavfi/avf_concat: return immediately after requesting a frame on input. · 79c1be12
  Nicolas George authored Oct 23, 2015
  
  79c1be12
- lavfi: remove astreamsync. · d92e0848
  Nicolas George authored Oct 24, 2015
```
It was only useful for very specific testing purposes
and appears to be currently partially broken.
```
  d92e0848