Commits · 4cf96c56420b89dc7145f906ac2bc67c880998ea · Linshizhi / ffmpeg.wasm-core

13 Dec, 2016 2 commits

swresample/resample: remove swri_resample function · 2b0112d4

Muhammad Faiz authored 8 years ago

integrate it inside multiple_resample
allow some calculations to be performed outside loop
Suggested-by: Michael Niedermayer <michael@niedermayer.cc>
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

2b0112d4

swresample/resample: do not allow negative dst_size return value · 6a8c0d83

Muhammad Faiz authored 8 years ago

This should fix Ticket6012
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

6a8c0d83

10 Dec, 2016 3 commits

swresample/resample_template: Add filter values in parallel · 65e33d8e

Michael Niedermayer authored 8 years ago

This is faster 2871 -> 2189  cycles for int16 matrixbench -> 23456hz
Fixes a integer overflow in a artificial corner case
Fixes part of 668007-media
Found-by: Matt Wolenetz <wolenetz@google.com>
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>

65e33d8e

swresample/resample_template: Reorder operations to avoid one addition · 34db6507
Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
34db6507

swresample/swresample: Check count before memcpy() · b3928a1c

Michael Niedermayer authored 8 years ago

Fixes undefined operation
Fixes part of 668007-media
Found-by: Matt Wolenetz <wolenetz@google.com>
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>

b3928a1c

03 Dec, 2016 1 commit
- swresample/resample: do not rebuild filter when sample_delta is zero · 01ebb57c
  Muhammad Faiz authored 8 years ago
```
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>
```
  01ebb57c
25 Nov, 2016 1 commit

swresample/soxr: fix invalid use of linear_interp · da34e4e1

Muhammad Faiz authored 8 years ago

give very bad quality for soxr resampler.
linear_interp is intended for  using linear interpolation
between filter bank so quality will be better.

i guess this is misunderstood as 'do not use filter bank,
but directly interpolate linearly between samples'.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

da34e4e1

24 Nov, 2016 1 commit

swresample/resample: optimize exact_rational=on:linear_interp=on case · 06f94149

Muhammad Faiz authored 8 years ago

separate dsp.resample to dsp.resample_common and dsp.resample_linear
and choose to call faster resample_common even when linear_interp=on
when c->frac and c->dst_incr_mod are both zero

speed up resampling when exact_rational and linear_interp are both
enabled because exact_rational force c->frac and c->dst_incr_mod to
be zero when soft compensation does not happen

benchmark on exact_rational=on:linear_interp=on
        old     new
real    8.432s  5.097s
user    7.679s  4.989s
sys     0.125s  0.107s
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

06f94149

26 Oct, 2016 3 commits
- Bump minor versions after 3.2 branchpoint to seperate release · 1609935b
  Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  1609935b
- Bump minor versions for 3.2 · 3f302520
  Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  3f302520
- swresample/rematrix: Fix float part of swr_set_matrix() · 9445e7e6
  Vodyannikov Aleksandr authored 8 years ago
```
Fixes Ticket #5897.
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  9445e7e6
18 Oct, 2016 1 commit
- swresample/resample: fix return value of build_filter · acd74f92
  Muhammad Faiz authored 8 years ago
```
return AVERROR code on error
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>
```
  acd74f92
27 Sep, 2016 3 commits
- swr: Update version & APIChanges for swr_build_matrix() · fd902510
  Michael Niedermayer authored 8 years ago
```
Found-by: James Almer <jamrial@gmail.com>
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  fd902510
- swresample: Add swr_build_matrix() · 23c0779c
  Michael Niedermayer authored 8 years ago
```
API and Doxy documentation is taken from avresample_build_matrix()
Fixes: Ticket5780
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  23c0779c
- swresample: Use double and float for matrixes for best quality and speed · 740f5105
  Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  740f5105
18 Aug, 2016 3 commits
- swresample: add int64 sample format · 9876d8fc
  Paul B Mahol authored 8 years ago
  
  9876d8fc
- swresample: Skip over dither steps if dithering scale is 0 · 30b2611e
  Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  30b2611e
- swresample: move dither init up · 946acacd
  Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  946acacd
03 Aug, 2016 1 commit
- doxygen: Standardize root-level modules · 58c7bf78
  Timothy Gu authored 8 years ago
  
  58c7bf78
22 Jun, 2016 1 commit
- swr: fix time.h include · a9eda4b2
  Clément Bœsch authored 8 years ago
  
  a9eda4b2
20 Jun, 2016 1 commit

swresample/x86: add support for exact_rational · 6031e5d1

Muhammad Faiz authored 8 years ago

phase_shift and phase_mask is removed
generally exact_rational=on is faster than exact_rational=off
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

6031e5d1

17 Jun, 2016 2 commits

swresample/resample: do not increase phase_count on exact_rational · 7f1b503e

Muhammad Faiz authored 8 years ago

high phase_count is only useful when dst_incr_mod is non zero
in other word, it is only useful on soft compensation

on init, it will build filter with low phase_count
but when soft compensation is enabled, rebuild filter
with high phase_count

this approach saves lots of memory
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

7f1b503e

swresample/resample: add support for odd phase_count · ee575acb

Muhammad Faiz authored 8 years ago

because exact_rational does not guarantee
that phase_count is even
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

ee575acb

13 Jun, 2016 1 commit

swresample: add exact_rational option · b8c6e5a6

Muhammad Faiz authored 8 years ago

give high quality resampling
as good as with linear_interp=on
as fast as without linear_interp=on
tested visually with ffplay
ffplay -f lavfi "aevalsrc='sin(10000*t*t)', aresample=osr=48000, showcqt=gamma=5"
ffplay -f lavfi "aevalsrc='sin(10000*t*t)', aresample=osr=48000:linear_interp=on, showcqt=gamma=5"
ffplay -f lavfi "aevalsrc='sin(10000*t*t)', aresample=osr=48000:exact_rational=on, showcqt=gamma=5"

slightly speed improvement
for fair comparison with -cpuflags 0
audio.wav is ~ 1 hour 44100 stereo 16bit wav file
ffmpeg -i audio.wav -af aresample=osr=48000 -f null -
        old         new
real    13.498s     13.121s
user    13.364s     12.987s
sys      0.131s      0.129s

linear_interp=on
        old         new
real    23.035s     23.050s
user    22.907s     22.917s
sys      0.119s     0.125s

exact_rational=on
real    12.418s
user    12.298s
sys      0.114s

possibility to decrease memory usage if soft compensation is ignored
Signed-off-by: Muhammad Faiz <mfcc64@gmail.com>

b8c6e5a6

16 May, 2016 1 commit
- swresample/resample: Fix division by 0 with tap_count=1 · feeb3a92
  Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  feeb3a92
15 May, 2016 2 commits

swresample/rematrix: Use clipping s16 rematrixing if overflows are possible · 2f76157e
Michael Niedermayer authored 8 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
2f76157e

swresample/rematrix: Use error diffusion to avoid error in the DC component of the matrix · 7fe81bc4

Michael Niedermayer authored 8 years ago

This fixes the sum of the integer coefficients ending up summing to a value
larger than the value representing unity.

This issue occurs with qN0.dts when converting to stereo
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>

7fe81bc4

13 May, 2016 1 commit
- swresample/arm: add ff_resample_common_apply_filter_{x4,x8}_{float,s16}_neon · f6265a5c
  Matthieu Bouron authored 8 years ago
  
  f6265a5c
22 Mar, 2016 1 commit
- swresample/swresample: Remove "less than" comparissions of enums · 914ad90e
  Michael Niedermayer authored 8 years ago
```
Found-by: wm4
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  914ad90e
14 Feb, 2016 1 commit

x86: use the new helper macros where useful · 70d685a7

James Almer authored 8 years ago

Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: James Almer <jamrial@gmail.com>

70d685a7

31 Jan, 2016 1 commit
- all: Make header guard names consistent · 180f9a09
  Timothy Gu authored 8 years ago
  
  180f9a09
24 Dec, 2015 1 commit

swr/resample: use av_clip_int16 instead of av_clip · 26937fb4

Ganesh Ajjanagadde authored 9 years ago

Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

26937fb4

04 Dec, 2015 1 commit
- swresample: use AV_OPT_TYPE_BOOL for linear_interp and cheby options · c1f114a8
  Clément Bœsch authored 9 years ago
  
  c1f114a8
15 Nov, 2015 1 commit

swresample/resample: remove redundant L for floating literal · 0bd0af6e

Ganesh Ajjanagadde authored 9 years ago

It is inherently double precision, and 1.0 is perfectly represented
anyway.
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

0bd0af6e

11 Nov, 2015 1 commit
- swresample/resample: increase precision for compensation · 351e625d
  Michael Niedermayer authored 9 years ago
```
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
```
  351e625d
09 Nov, 2015 4 commits

swresample/resample: speed up Blackman Nuttall filter · cf491a92

Ganesh Ajjanagadde authored 9 years ago

This may be a slightly surprising optimization, but is actually based on
an understanding of how math libraries compute trigonometric functions.
Explanation is given here so that future development uses libm more effectively
across the codebase.

All libm's essentially compute transcendental functions via some kind of
polynomial approximation, be it Taylor-Maclaurin or Chebyshev.
Correction terms are added via polynomial correction factors when needed
to squeeze out the last bits of accuracy. Lookup tables are also
inserted strategically.

In the case of trigonometric functions, periodicity is exploited via
first doing a range reduction to an interval around zero, and then using
some polynomial approximation.

This range reduction is the most natural way of doing things - else one
would need polynomials for ranges in different periods which makes no
sense whatsoever.

To avoid the need for the range reduction, it is helpful to feed in
arguments as close to the origin as possible for the trigonometric
functions. In fact, this also makes sense from an accuracy point of view:
IEEE floating point has far more resolution for small numbers than big ones.

This patch does this for the Blackman-Nuttall filter, and yields a
non-negligible speedup.

Sample benchmark (x86-64, Haswell, GNU/Linux)
test: fate-swr-resample-dblp-2626-44100
old:
18893514 decicycles in build_filter (loop 1000),     256 runs,      0 skips
18599863 decicycles in build_filter (loop 1000),     512 runs,      0 skips
18445574 decicycles in build_filter (loop 1000),    1000 runs,     24 skips

new:
16290697 decicycles in build_filter (loop 1000),     256 runs,      0 skips
16267172 decicycles in build_filter (loop 1000),     512 runs,      0 skips
16251105 decicycles in build_filter (loop 1000),    1000 runs,     24 skips
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

cf491a92

swresample/resample: speed up upsampling by precomputing sines · b87ca4bf

Ganesh Ajjanagadde authored 9 years ago

When upsampling, factor is set to 1 and sines need to be evaluated only
once for each phase, and the complexity should not depend on the number
of filter taps. This does the desired precomputation, yielding
significant speedups. Hard guarantees on the gain are not possible, but gains
themselves are obvious and are illustrated below.

Sample benchmark (x86-64, Haswell, GNU/Linux)
test: fate-swr-resample-dblp-2626-44100
old:
29161085 decicycles in build_filter (loop 1000),     256 runs,      0 skips
28821467 decicycles in build_filter (loop 1000),     512 runs,      0 skips
28668201 decicycles in build_filter (loop 1000),    1000 runs,     24 skips

new:
14351936 decicycles in build_filter (loop 1000),     256 runs,      0 skips
14306652 decicycles in build_filter (loop 1000),     512 runs,      0 skips
14299923 decicycles in build_filter (loop 1000),    1000 runs,     24 skips

Note that this does not statically allocate the sin lookup table. This
may be done for the default 1024 phases, yielding a 512*8 = 4kB array
which should be small enough.
This should yield a small improvement. Nevertheless, this is separate from
this patch, is more ambiguous due to the binary increase, and requires a
lut to be generated offline.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

b87ca4bf

swresample/resample: improve bessel function accuracy and speed · a5202bc9

Ganesh Ajjanagadde authored 9 years ago

This improves accuracy for the bessel function at large arguments, and this in turn
should improve the quality of the Kaiser window. It also improves the
performance of the bessel function and hence build_filter by ~ 20%.
Details are given below.

Algorithm: taken from the Boost project, who have done a detailed
investigation of the accuracy of their method, as compared with e.g the
GNU Scientific Library (GSL):
http://www.boost.org/doc/libs/1_52_0/libs/math/doc/sf_and_dist/html/math_toolkit/special/bessel/mbessel.html.
Boost source code (also cited and licensed in the code):
https://searchcode.com/codesearch/view/14918379/.

Accuracy: sample values may be obtained as follows. i0 denotes the old bessel code,
i0_boost the approach here, and i0_real an arbitrary precision result (truncated) from Wolfram Alpha:
type "bessel i0(6.0)" to reproduce. These are evaluation points that occur for
the default kaiser_beta = 9.

Some illustrations:
bessel(8.0)
i0      (8.000000) = 427.564115721804739678191254
i0_boost(8.000000) = 427.564115721804796521610115
i0_real (8.000000) = 427.564115721804785177396791

bessel(6.0)
i0      (6.000000) = 67.234406976477956163762428
i0_boost(6.000000) = 67.234406976477970374617144
i0_real (6.000000) = 67.234406976477975326188025

Reason for accuracy: Main accuracy benefits come at larger bessel arguments, where the
Taylor-Maclaurin method is not that good: 23+ iterations
(at large arguments, since the series is about 0) can cause
significant floating point error accumulation.

Benchmarks: Obtained on x86-64, Haswell, GNU/Linux via a loop calling
build_filter 1000 times:
test: fate-swr-resample-dblp-44100-2626

new:
995894468 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1029719302 decicycles in build_filter(loop 1000),     512 runs,      0 skips
984101131 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

old:
1250020763 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1246353282 decicycles in build_filter(loop 1000),     512 runs,      0 skips
1220017565 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

A further ~ 5% may be squeezed by enabling -ftree-vectorize. However,
this is a separate issue from this patch.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

a5202bc9

swresample: allow double precision beta value for the Kaiser window · 1bed09a3

Ganesh Ajjanagadde authored 9 years ago

Kaiser windows inherently don't require beta to be an integer. This was
an arbitrary restriction. Moreover, soxr does not require it, and in
fact often estimates beta to a non-integral value.

Thus, this patch allows greater flexibility for swresample clients.
Micro version is updated.
Reviewed-by: Derek Buitenhuis <derek.buitenhuis@gmail.com>
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

1bed09a3

06 Nov, 2015 1 commit

swresample/resample: speed up build_filter for Blackman-Nuttall filter · c8780822

Ganesh Ajjanagadde authored 9 years ago

This uses the trigonometric double and triple angle formulae to avoid
repeated (expensive) evaluation of libc's cos().

Sample benchmark (x86-64, Haswell, GNU/Linux)
test: fate-swr-resample-dblp-44100-2626
old:
1104466600 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1096765286 decicycles in build_filter(loop 1000),     512 runs,      0 skips
1070479590 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

new:
588861423 decicycles in build_filter(loop 1000),     256 runs,      0 skips
591262754 decicycles in build_filter(loop 1000),     512 runs,      0 skips
577355145 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

This results in small differences with the old expression:
difference (worst case on [0, 2*M_PI]), argmax 0.008:
max diff (relative): 0.000000000000157289807188
blackman_old(0.008): 0.000363951585488813192382
blackman_new(0.008): 0.000363951585488755946507

These are judged to be insignificant for the performance gain. PSNR to
reference file is unchanged up to second decimal point for instance.
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

c8780822