commit 00c2b78

uint  ·  2026-07-29 13:38:17 +0000 UTC
parent af0e98b
vendor c89 translation of stb_image
1 files changed,  +8244, -0
+8244, -0
   1@@ -0,0 +1,8244 @@
   2+/* stb_image - v2.30 - public domain image loader - http://nothings.org/stb
   3+                                  no warranty implied; use at your own risk
   4+
   5+   Do this:
   6+      #define STB_IMAGE_IMPLEMENTATION
   7+   before you include this file in *one* C or C++ file to create the implementation.
   8+
   9+   // i.e. it should look like this:
  10+   #include ...
  11+   #include ...
  12+   #include ...
  13+   #define STB_IMAGE_IMPLEMENTATION
  14+   #include "stb_image.h"
  15+
  16+   You can #define STBI_ASSERT(x) before the #include to avoid using assert.h.
  17+   And #define STBI_MALLOC, STBI_REALLOC, and STBI_FREE to avoid using malloc,realloc,free
  18+
  19+
  20+   QUICK NOTES:
  21+      Primarily of interest to game developers and other people who can
  22+          avoid problematic images and only need the trivial interface
  23+
  24+      JPEG baseline & progressive (12 bpc/arithmetic not supported, same as stock IJG lib)
  25+      PNG 1/2/4/8/16-bit-per-channel
  26+
  27+      TGA (not sure what subset, if a subset)
  28+      BMP non-1bpp, non-RLE
  29+      PSD (composited view only, no extra channels, 8/16 bit-per-channel)
  30+
  31+      GIF (*comp always reports as 4-channel)
  32+      HDR (radiance rgbE format)
  33+      PIC (Softimage PIC)
  34+      PNM (PPM and PGM binary only)
  35+
  36+      Animated GIF still needs a proper API, but here's one way to do it:
  37+          http://gist.github.com/urraka/685d9a6340b26b830d49
  38+
  39+      - decode from memory or through FILE (define STBI_NO_STDIO to remove code)
  40+      - decode from arbitrary I/O callbacks
  41+      - SIMD acceleration on x86/x64 (SSE2) and ARM (NEON)
  42+
  43+   Full documentation under "DOCUMENTATION" below.
  44+
  45+
  46+LICENSE
  47+
  48+  See end of file for license information.
  49+
  50+RECENT REVISION HISTORY:
  51+
  52+      2.30  (2024-05-31) avoid erroneous gcc warning
  53+      2.29  (2023-05-xx) optimizations
  54+      2.28  (2023-01-29) many error fixes, security errors, just tons of stuff
  55+      2.27  (2021-07-11) document stbi_info better, 16-bit PNM support, bug fixes
  56+      2.26  (2020-07-13) many minor fixes
  57+      2.25  (2020-02-02) fix warnings
  58+      2.24  (2020-02-02) fix warnings; thread-local failure_reason and flip_vertically
  59+      2.23  (2019-08-11) fix clang static analysis warning
  60+      2.22  (2019-03-04) gif fixes, fix warnings
  61+      2.21  (2019-02-25) fix typo in comment
  62+      2.20  (2019-02-07) support utf8 filenames in Windows; fix warnings and platform ifdefs
  63+      2.19  (2018-02-11) fix warning
  64+      2.18  (2018-01-30) fix warnings
  65+      2.17  (2018-01-29) bugfix, 1-bit BMP, 16-bitness query, fix warnings
  66+      2.16  (2017-07-23) all functions have 16-bit variants; optimizations; bugfixes
  67+      2.15  (2017-03-18) fix png-1,2,4; all Imagenet JPGs; no runtime SSE detection on GCC
  68+      2.14  (2017-03-03) remove deprecated STBI_JPEG_OLD; fixes for Imagenet JPGs
  69+      2.13  (2016-12-04) experimental 16-bit API, only for PNG so far; fixes
  70+      2.12  (2016-04-02) fix typo in 2.11 PSD fix that caused crashes
  71+      2.11  (2016-04-02) 16-bit PNGS; enable SSE2 in non-gcc x64
  72+                         RGB-format JPEG; remove white matting in PSD;
  73+                         allocate large structures on the stack;
  74+                         correct channel count for PNG & BMP
  75+      2.10  (2016-01-22) avoid warning introduced in 2.09
  76+      2.09  (2016-01-16) 16-bit TGA; comments in PNM files; STBI_REALLOC_SIZED
  77+
  78+   See end of file for full revision history.
  79+
  80+
  81+ ============================    Contributors    =========================
  82+
  83+ Image formats                          Extensions, features
  84+    Sean Barrett (jpeg, png, bmp)          Jetro Lauha (stbi_info)
  85+    Nicolas Schulz (hdr, psd)              Martin "SpartanJ" Golini (stbi_info)
  86+    Jonathan Dummer (tga)                  James "moose2000" Brown (iPhone PNG)
  87+    Jean-Marc Lienher (gif)                Ben "Disch" Wenger (io callbacks)
  88+    Tom Seddon (pic)                       Omar Cornut (1/2/4-bit PNG)
  89+    Thatcher Ulrich (psd)                  Nicolas Guillemot (vertical flip)
  90+    Ken Miller (pgm, ppm)                  Richard Mitton (16-bit PSD)
  91+    github:urraka (animated gif)           Junggon Kim (PNM comments)
  92+    Christopher Forseth (animated gif)     Daniel Gibson (16-bit TGA)
  93+                                           socks-the-fox (16-bit PNG)
  94+                                           Jeremy Sawicki (handle all ImageNet JPGs)
  95+ Optimizations & bugfixes                  Mikhail Morozov (1-bit BMP)
  96+    Fabian "ryg" Giesen                    Anael Seghezzi (is-16-bit query)
  97+    Arseny Kapoulkine                      Simon Breuss (16-bit PNM)
  98+    John-Mark Allen
  99+    Carmelo J Fdez-Aguera
 100+
 101+ Bug & warning fixes
 102+    Marc LeBlanc            David Woo          Guillaume George     Martins Mozeiko
 103+    Christpher Lloyd        Jerry Jansson      Joseph Thomson       Blazej Dariusz Roszkowski
 104+    Phil Jordan                                Dave Moore           Roy Eltham
 105+    Hayaki Saito            Nathan Reed        Won Chun
 106+    Luke Graham             Johan Duparc       Nick Verigakis       the Horde3D community
 107+    Thomas Ruf              Ronny Chevalier                         github:rlyeh
 108+    Janez Zemva             John Bartholomew   Michal Cichon        github:romigrou
 109+    Jonathan Blow           Ken Hamada         Tero Hanninen        github:svdijk
 110+    Eugene Golushkov        Laurent Gomila     Cort Stratton        github:snagar
 111+    Aruelien Pocheville     Sergio Gonzalez    Thibault Reuille     github:Zelex
 112+    Cass Everitt            Ryamond Barbiero                        github:grim210
 113+    Paul Du Bois            Engin Manap        Aldo Culquicondor    github:sammyhw
 114+    Philipp Wiesemann       Dale Weiler        Oriol Ferrer Mesia   github:phprus
 115+    Josh Tobin              Neil Bickford      Matthew Gregan       github:poppolopoppo
 116+    Julian Raschke          Gregory Mullen     Christian Floisand   github:darealshinji
 117+    Baldur Karlsson         Kevin Schmidt      JR Smith             github:Michaelangel007
 118+                            Brad Weinberger    Matvey Cherevko      github:mosra
 119+    Luca Sas                Alexander Veselov  Zack Middleton       [reserved]
 120+    Ryan C. Gordon          [reserved]                              [reserved]
 121+                     DO NOT ADD YOUR NAME HERE
 122+
 123+                     Jacko Dirks
 124+
 125+  To add your name to the credits, pick a random blank space in the middle and fill it.
 126+  80% of merge conflicts on stb PRs are due to people adding their name at the end
 127+  of the credits.
 128+*/
 129+
 130+#ifndef STBI_INCLUDE_STB_IMAGE_H
 131+#define STBI_INCLUDE_STB_IMAGE_H
 132+
 133+/*
 134+ * DOCUMENTATION
 135+ *
 136+ * Limitations:
 137+ *    - no 12-bit-per-channel JPEG
 138+ *    - no JPEGs with arithmetic coding
 139+ *    - GIF always returns *comp=4
 140+ *
 141+ * Basic usage (see HDR discussion below for HDR usage):
 142+ *    int x,y,n;
 143+ *    unsigned char *data = stbi_load(filename, &x, &y, &n, 0);
 144+ *    // ... process data if not NULL ...
 145+ *    // ... x = width, y = height, n = # 8-bit components per pixel ...
 146+ *    // ... replace '0' with '1'..'4' to force that many components per pixel
 147+ *    // ... but 'n' will always be the number that it would have been if you said 0
 148+ *    stbi_image_free(data);
 149+ *
 150+ * Standard parameters:
 151+ *    int *x                 -- outputs image width in pixels
 152+ *    int *y                 -- outputs image height in pixels
 153+ *    int *channels_in_file  -- outputs # of image components in image file
 154+ *    int desired_channels   -- if non-zero, # of image components requested in result
 155+ *
 156+ * The return value from an image loader is an 'unsigned char *' which points
 157+ * to the pixel data, or NULL on an allocation failure or if the image is
 158+ * corrupt or invalid. The pixel data consists of *y scanlines of *x pixels,
 159+ * with each pixel consisting of N interleaved 8-bit components; the first
 160+ * pixel pointed to is top-left-most in the image. There is no padding between
 161+ * image scanlines or between pixels, regardless of format. The number of
 162+ * components N is 'desired_channels' if desired_channels is non-zero, or
 163+ * *channels_in_file otherwise. If desired_channels is non-zero,
 164+ * *channels_in_file has the number of components that _would_ have been
 165+ * output otherwise. E.g. if you set desired_channels to 4, you will always
 166+ * get RGBA output, but you can check *channels_in_file to see if it's trivially
 167+ * opaque because e.g. there were only 3 channels in the source image.
 168+ *
 169+ * An output image with N components has the following components interleaved
 170+ * in this order in each pixel:
 171+ *
 172+ *     N=#comp     components
 173+ *       1           grey
 174+ *       2           grey, alpha
 175+ *       3           red, green, blue
 176+ *       4           red, green, blue, alpha
 177+ *
 178+ * If image loading fails for any reason, the return value will be NULL,
 179+ * and *x, *y, *channels_in_file will be unchanged. The function
 180+ * stbi_failure_reason() can be queried for an extremely brief, end-user
 181+ * unfriendly explanation of why the load failed. Define STBI_NO_FAILURE_STRINGS
 182+ * to avoid compiling these strings at all, and STBI_FAILURE_USERMSG to get slightly
 183+ * more user-friendly ones.
 184+ *
 185+ * Paletted PNG, BMP, GIF, and PIC images are automatically depalettized.
 186+ *
 187+ * To query the width, height and component count of an image without having to
 188+ * decode the full file, you can use the stbi_info family of functions:
 189+ *
 190+ *   int x,y,n,ok;
 191+ *   ok = stbi_info(filename, &x, &y, &n);
 192+ *   // returns ok=1 and sets x, y, n if image is a supported format,
 193+ *   // 0 otherwise.
 194+ *
 195+ * Note that stb_image pervasively uses ints in its public API for sizes,
 196+ * including sizes of memory buffers. This is now part of the API and thus
 197+ * hard to change without causing breakage. As a result, the various image
 198+ * loaders all have certain limits on image size; these differ somewhat
 199+ * by format but generally boil down to either just under 2GB or just under
 200+ * 1GB. When the decoded image would be larger than this, stb_image decoding
 201+ * will fail.
 202+ *
 203+ * Additionally, stb_image will reject image files that have any of their
 204+ * dimensions set to a larger value than the configurable STBI_MAX_DIMENSIONS,
 205+ * which defaults to 2**24 = 16777216 pixels. Due to the above memory limit,
 206+ * the only way to have an image with such dimensions load correctly
 207+ * is for it to have a rather extreme aspect ratio. Either way, the
 208+ * assumption here is that such larger images are likely to be malformed
 209+ * or malicious. If you do need to load an image with individual dimensions
 210+ * larger than that, and it still fits in the overall size limit, you can
 211+ * #define STBI_MAX_DIMENSIONS on your own to be something larger.
 212+ *
 213+ */
 214+/*===========================================================================*/
 215+/*
 216+ *
 217+ * UNICODE:
 218+ *
 219+ *   If compiling for Windows and you wish to use Unicode filenames, compile
 220+ *   with
 221+ *       #define STBI_WINDOWS_UTF8
 222+ *   and pass utf8-encoded filenames. Call stbi_convert_wchar_to_utf8 to convert
 223+ *   Windows wchar_t filenames to utf8.
 224+ *
 225+ */
 226+/*===========================================================================*/
 227+/*
 228+ *
 229+ * Philosophy
 230+ *
 231+ * stb libraries are designed with the following priorities:
 232+ *
 233+ *    1. easy to use
 234+ *    2. easy to maintain
 235+ *    3. good performance
 236+ *
 237+ * Sometimes I let "good performance" creep up in priority over "easy to maintain",
 238+ * and for best performance I may provide less-easy-to-use APIs that give higher
 239+ * performance, in addition to the easy-to-use ones. Nevertheless, it's important
 240+ * to keep in mind that from the standpoint of you, a client of this library,
 241+ * all you care about is #1 and #3, and stb libraries DO NOT emphasize #3 above all.
 242+ *
 243+ * Some secondary priorities arise directly from the first two, some of which
 244+ * provide more explicit reasons why performance can't be emphasized.
 245+ *
 246+ *    - Portable ("ease of use")
 247+ *    - Small source code footprint ("easy to maintain")
 248+ *    - No dependencies ("ease of use")
 249+ *
 250+ */
 251+/*===========================================================================*/
 252+/*
 253+ *
 254+ * I/O callbacks
 255+ *
 256+ * I/O callbacks allow you to read from arbitrary sources, like packaged
 257+ * files or some other source. Data read from callbacks are processed
 258+ * through a small internal buffer (currently 128 bytes) to try to reduce
 259+ * overhead.
 260+ *
 261+ * The three functions you must define are "read" (reads some bytes of data),
 262+ * "skip" (skips some bytes of data), "eof" (reports if the stream is at the end).
 263+ *
 264+ */
 265+/*===========================================================================*/
 266+/*
 267+ *
 268+ * SIMD support
 269+ *
 270+ * The JPEG decoder will try to automatically use SIMD kernels on x86 when
 271+ * supported by the compiler. For ARM Neon support, you must explicitly
 272+ * request it.
 273+ *
 274+ * (The old do-it-yourself SIMD API is no longer supported in the current
 275+ * code.)
 276+ *
 277+ * On x86, SSE2 will automatically be used when available based on a run-time
 278+ * test; if not, the generic C versions are used as a fall-back. On ARM targets,
 279+ * the typical path is to have separate builds for NEON and non-NEON devices
 280+ * (at least this is true for iOS and Android). Therefore, the NEON support is
 281+ * toggled by a build flag: define STBI_NEON to get NEON loops.
 282+ *
 283+ * If for some reason you do not want to use any of SIMD code, or if
 284+ * you have issues compiling it, you can disable it entirely by
 285+ * defining STBI_NO_SIMD.
 286+ *
 287+ */
 288+/*===========================================================================*/
 289+/*
 290+ *
 291+ * HDR image support   (disable by defining STBI_NO_HDR)
 292+ *
 293+ * stb_image supports loading HDR images in general, and currently the Radiance
 294+ * .HDR file format specifically. You can still load any file through the existing
 295+ * interface; if you attempt to load an HDR file, it will be automatically remapped
 296+ * to LDR, assuming gamma 2.2 and an arbitrary scale factor defaulting to 1;
 297+ * both of these constants can be reconfigured through this interface:
 298+ *
 299+ *     stbi_hdr_to_ldr_gamma(2.2f);
 300+ *     stbi_hdr_to_ldr_scale(1.0f);
 301+ *
 302+ * (note, do not use _inverse_ constants; stbi_image will invert them
 303+ * appropriately).
 304+ *
 305+ * Additionally, there is a new, parallel interface for loading files as
 306+ * (linear) floats to preserve the full dynamic range:
 307+ *
 308+ *    float *data = stbi_loadf(filename, &x, &y, &n, 0);
 309+ *
 310+ * If you load LDR images through this interface, those images will
 311+ * be promoted to floating point values, run through the inverse of
 312+ * constants corresponding to the above:
 313+ *
 314+ *     stbi_ldr_to_hdr_scale(1.0f);
 315+ *     stbi_ldr_to_hdr_gamma(2.2f);
 316+ *
 317+ * Finally, given a filename (or an open file or memory block--see header
 318+ * file for details) containing image data, you can query for the "most
 319+ * appropriate" interface to use (that is, whether the image is HDR or
 320+ * not), using:
 321+ *
 322+ *     stbi_is_hdr(char *filename);
 323+ *
 324+ */
 325+/*===========================================================================*/
 326+/*
 327+ *
 328+ * iPhone PNG support:
 329+ *
 330+ * We optionally support converting iPhone-formatted PNGs (which store
 331+ * premultiplied BGRA) back to RGB, even though they're internally encoded
 332+ * differently. To enable this conversion, call
 333+ * stbi_convert_iphone_png_to_rgb(1).
 334+ *
 335+ * Call stbi_set_unpremultiply_on_load(1) as well to force a divide per
 336+ * pixel to remove any premultiplied alpha *only* if the image file explicitly
 337+ * says there's premultiplied data (currently only happens in iPhone images,
 338+ * and only if iPhone convert-to-rgb processing is on).
 339+ *
 340+ */
 341+/*===========================================================================*/
 342+/*
 343+ *
 344+ * ADDITIONAL CONFIGURATION
 345+ *
 346+ *  - You can suppress implementation of any of the decoders to reduce
 347+ *    your code footprint by #defining one or more of the following
 348+ *    symbols before creating the implementation.
 349+ *
 350+ *        STBI_NO_JPEG
 351+ *        STBI_NO_PNG
 352+ *        STBI_NO_BMP
 353+ *        STBI_NO_PSD
 354+ *        STBI_NO_TGA
 355+ *        STBI_NO_GIF
 356+ *        STBI_NO_HDR
 357+ *        STBI_NO_PIC
 358+ *        STBI_NO_PNM   (.ppm and .pgm)
 359+ *
 360+ *  - You can request *only* certain decoders and suppress all other ones
 361+ *    (this will be more forward-compatible, as addition of new decoders
 362+ *    doesn't require you to disable them explicitly):
 363+ *
 364+ *        STBI_ONLY_JPEG
 365+ *        STBI_ONLY_PNG
 366+ *        STBI_ONLY_BMP
 367+ *        STBI_ONLY_PSD
 368+ *        STBI_ONLY_TGA
 369+ *        STBI_ONLY_GIF
 370+ *        STBI_ONLY_HDR
 371+ *        STBI_ONLY_PIC
 372+ *        STBI_ONLY_PNM   (.ppm and .pgm)
 373+ *
 374+ *   - If you use STBI_NO_PNG (or _ONLY_ without PNG), and you still
 375+ *     want the zlib decoder to be available, #define STBI_SUPPORT_ZLIB
 376+ *
 377+ *  - If you define STBI_MAX_DIMENSIONS, stb_image will reject images greater
 378+ *    than that size (in either width or height) without further processing.
 379+ *    This is to let programs in the wild set an upper bound to prevent
 380+ *    denial-of-service attacks on untrusted data, as one could generate a
 381+ *    valid image of gigantic dimensions and force stb_image to allocate a
 382+ *    huge block of memory and spend disproportionate time decoding it. By
 383+ *    default this is set to (1 << 24), which is 16777216, but that's still
 384+ *    very big.
 385+ */
 386+
 387+#ifndef STBI_NO_STDIO
 388+#include <stdio.h>
 389+#endif /* STBI_NO_STDIO */
 390+
 391+#define STBI_VERSION 1
 392+
 393+enum
 394+{
 395+   STBI_default = 0, /* only used for desired_channels */
 396+
 397+   STBI_grey       = 1,
 398+   STBI_grey_alpha = 2,
 399+   STBI_rgb        = 3,
 400+   STBI_rgb_alpha  = 4
 401+};
 402+
 403+#include <stdlib.h>
 404+typedef unsigned char stbi_uc;
 405+typedef unsigned short stbi_us;
 406+
 407+#ifdef __cplusplus
 408+extern "C" {
 409+#endif
 410+
 411+#ifndef STBIDEF
 412+#ifdef STB_IMAGE_STATIC
 413+#define STBIDEF static
 414+#else
 415+#define STBIDEF extern
 416+#endif
 417+#endif
 418+
 419+/*////////////////////////////////////////////////////////////////////////////*/
 420+/*
 421+ *
 422+ * PRIMARY API - works on images of any type
 423+ *
 424+ */
 425+
 426+/*
 427+ *
 428+ * load image by filename, open file, or memory buffer
 429+ *
 430+ */
 431+
 432+typedef struct
 433+{
 434+   int      (*read)  (void *user,char *data,int size); /* fill 'data' with 'size' bytes.  return number of bytes actually read */
 435+   void     (*skip)  (void *user,int n); /* skip the next 'n' bytes, or 'unget' the last -n bytes if negative */
 436+   int      (*eof)   (void *user); /* returns nonzero if we are at end of file/data */
 437+} stbi_io_callbacks;
 438+
 439+/*//////////////////////////////////*/
 440+/*
 441+ *
 442+ * 8-bits-per-channel interface
 443+ *
 444+ */
 445+
 446+STBIDEF stbi_uc *stbi_load_from_memory   (stbi_uc           const *buffer, int len   , int *x, int *y, int *channels_in_file, int desired_channels);
 447+STBIDEF stbi_uc *stbi_load_from_callbacks(stbi_io_callbacks const *clbk  , void *user, int *x, int *y, int *channels_in_file, int desired_channels);
 448+
 449+#ifndef STBI_NO_STDIO
 450+STBIDEF stbi_uc *stbi_load            (char const *filename, int *x, int *y, int *channels_in_file, int desired_channels);
 451+STBIDEF stbi_uc *stbi_load_from_file  (FILE *f, int *x, int *y, int *channels_in_file, int desired_channels);
 452+/* for stbi_load_from_file, file pointer is left pointing immediately after image */
 453+#endif
 454+
 455+#ifndef STBI_NO_GIF
 456+STBIDEF stbi_uc *stbi_load_gif_from_memory(stbi_uc const *buffer, int len, int **delays, int *x, int *y, int *z, int *comp, int req_comp);
 457+#endif
 458+
 459+#ifdef STBI_WINDOWS_UTF8
 460+STBIDEF int stbi_convert_wchar_to_utf8(char *buffer, size_t bufferlen, const wchar_t* input);
 461+#endif
 462+
 463+/*//////////////////////////////////*/
 464+/*
 465+ *
 466+ * 16-bits-per-channel interface
 467+ *
 468+ */
 469+
 470+STBIDEF stbi_us *stbi_load_16_from_memory   (stbi_uc const *buffer, int len, int *x, int *y, int *channels_in_file, int desired_channels);
 471+STBIDEF stbi_us *stbi_load_16_from_callbacks(stbi_io_callbacks const *clbk, void *user, int *x, int *y, int *channels_in_file, int desired_channels);
 472+
 473+#ifndef STBI_NO_STDIO
 474+STBIDEF stbi_us *stbi_load_16          (char const *filename, int *x, int *y, int *channels_in_file, int desired_channels);
 475+STBIDEF stbi_us *stbi_load_from_file_16(FILE *f, int *x, int *y, int *channels_in_file, int desired_channels);
 476+#endif
 477+
 478+/*//////////////////////////////////*/
 479+/*
 480+ *
 481+ * float-per-channel interface
 482+ *
 483+ */
 484+#ifndef STBI_NO_LINEAR
 485+   STBIDEF float *stbi_loadf_from_memory     (stbi_uc const *buffer, int len, int *x, int *y, int *channels_in_file, int desired_channels);
 486+   STBIDEF float *stbi_loadf_from_callbacks  (stbi_io_callbacks const *clbk, void *user, int *x, int *y,  int *channels_in_file, int desired_channels);
 487+
 488+   #ifndef STBI_NO_STDIO
 489+   STBIDEF float *stbi_loadf            (char const *filename, int *x, int *y, int *channels_in_file, int desired_channels);
 490+   STBIDEF float *stbi_loadf_from_file  (FILE *f, int *x, int *y, int *channels_in_file, int desired_channels);
 491+   #endif
 492+#endif
 493+
 494+#ifndef STBI_NO_HDR
 495+   STBIDEF void   stbi_hdr_to_ldr_gamma(float gamma);
 496+   STBIDEF void   stbi_hdr_to_ldr_scale(float scale);
 497+#endif /* STBI_NO_HDR */
 498+
 499+#ifndef STBI_NO_LINEAR
 500+   STBIDEF void   stbi_ldr_to_hdr_gamma(float gamma);
 501+   STBIDEF void   stbi_ldr_to_hdr_scale(float scale);
 502+#endif /* STBI_NO_LINEAR */
 503+
 504+/* stbi_is_hdr is always defined, but always returns false if STBI_NO_HDR */
 505+STBIDEF int    stbi_is_hdr_from_callbacks(stbi_io_callbacks const *clbk, void *user);
 506+STBIDEF int    stbi_is_hdr_from_memory(stbi_uc const *buffer, int len);
 507+#ifndef STBI_NO_STDIO
 508+STBIDEF int      stbi_is_hdr          (char const *filename);
 509+STBIDEF int      stbi_is_hdr_from_file(FILE *f);
 510+#endif /* STBI_NO_STDIO */
 511+
 512+
 513+/*
 514+ * get a VERY brief reason for failure
 515+ * on most compilers (and ALL modern mainstream compilers) this is threadsafe
 516+ */
 517+STBIDEF const char *stbi_failure_reason  (void);
 518+
 519+/* free the loaded image -- this is just free() */
 520+STBIDEF void     stbi_image_free      (void *retval_from_stbi_load);
 521+
 522+/* get image dimensions & components without fully decoding */
 523+STBIDEF int      stbi_info_from_memory(stbi_uc const *buffer, int len, int *x, int *y, int *comp);
 524+STBIDEF int      stbi_info_from_callbacks(stbi_io_callbacks const *clbk, void *user, int *x, int *y, int *comp);
 525+STBIDEF int      stbi_is_16_bit_from_memory(stbi_uc const *buffer, int len);
 526+STBIDEF int      stbi_is_16_bit_from_callbacks(stbi_io_callbacks const *clbk, void *user);
 527+
 528+#ifndef STBI_NO_STDIO
 529+STBIDEF int      stbi_info               (char const *filename,     int *x, int *y, int *comp);
 530+STBIDEF int      stbi_info_from_file     (FILE *f,                  int *x, int *y, int *comp);
 531+STBIDEF int      stbi_is_16_bit          (char const *filename);
 532+STBIDEF int      stbi_is_16_bit_from_file(FILE *f);
 533+#endif
 534+
 535+
 536+
 537+/*
 538+ * for image formats that explicitly notate that they have premultiplied alpha,
 539+ * we just return the colors as stored in the file. set this flag to force
 540+ * unpremultiplication. results are undefined if the unpremultiply overflow.
 541+ */
 542+STBIDEF void stbi_set_unpremultiply_on_load(int flag_true_if_should_unpremultiply);
 543+
 544+/*
 545+ * indicate whether we should process iphone images back to canonical format,
 546+ * or just pass them through "as-is"
 547+ */
 548+STBIDEF void stbi_convert_iphone_png_to_rgb(int flag_true_if_should_convert);
 549+
 550+/* flip the image vertically, so the first pixel in the output array is the bottom left */
 551+STBIDEF void stbi_set_flip_vertically_on_load(int flag_true_if_should_flip);
 552+
 553+/*
 554+ * as above, but only applies to images loaded on the thread that calls the function
 555+ * this function is only available if your compiler supports thread-local variables;
 556+ * calling it will fail to link if your compiler doesn't
 557+ */
 558+STBIDEF void stbi_set_unpremultiply_on_load_thread(int flag_true_if_should_unpremultiply);
 559+STBIDEF void stbi_convert_iphone_png_to_rgb_thread(int flag_true_if_should_convert);
 560+STBIDEF void stbi_set_flip_vertically_on_load_thread(int flag_true_if_should_flip);
 561+
 562+/* ZLIB client - used by PNG, available for other purposes */
 563+
 564+STBIDEF char *stbi_zlib_decode_malloc_guesssize(const char *buffer, int len, int initial_size, int *outlen);
 565+STBIDEF char *stbi_zlib_decode_malloc_guesssize_headerflag(const char *buffer, int len, int initial_size, int *outlen, int parse_header);
 566+STBIDEF char *stbi_zlib_decode_malloc(const char *buffer, int len, int *outlen);
 567+STBIDEF int   stbi_zlib_decode_buffer(char *obuffer, int olen, const char *ibuffer, int ilen);
 568+
 569+STBIDEF char *stbi_zlib_decode_noheader_malloc(const char *buffer, int len, int *outlen);
 570+STBIDEF int   stbi_zlib_decode_noheader_buffer(char *obuffer, int olen, const char *ibuffer, int ilen);
 571+
 572+
 573+#ifdef __cplusplus
 574+}
 575+#endif
 576+
 577+/*
 578+ *
 579+ *
 580+ * //   end header file   /////////////////////////////////////////////////////
 581+ */
 582+#endif /* STBI_INCLUDE_STB_IMAGE_H */
 583+
 584+#ifdef STB_IMAGE_IMPLEMENTATION
 585+
 586+#if defined(STBI_ONLY_JPEG) || defined(STBI_ONLY_PNG) || defined(STBI_ONLY_BMP) \
 587+  || defined(STBI_ONLY_TGA) || defined(STBI_ONLY_GIF) || defined(STBI_ONLY_PSD) \
 588+  || defined(STBI_ONLY_HDR) || defined(STBI_ONLY_PIC) || defined(STBI_ONLY_PNM) \
 589+  || defined(STBI_ONLY_ZLIB)
 590+   #ifndef STBI_ONLY_JPEG
 591+   #define STBI_NO_JPEG
 592+   #endif
 593+   #ifndef STBI_ONLY_PNG
 594+   #define STBI_NO_PNG
 595+   #endif
 596+   #ifndef STBI_ONLY_BMP
 597+   #define STBI_NO_BMP
 598+   #endif
 599+   #ifndef STBI_ONLY_PSD
 600+   #define STBI_NO_PSD
 601+   #endif
 602+   #ifndef STBI_ONLY_TGA
 603+   #define STBI_NO_TGA
 604+   #endif
 605+   #ifndef STBI_ONLY_GIF
 606+   #define STBI_NO_GIF
 607+   #endif
 608+   #ifndef STBI_ONLY_HDR
 609+   #define STBI_NO_HDR
 610+   #endif
 611+   #ifndef STBI_ONLY_PIC
 612+   #define STBI_NO_PIC
 613+   #endif
 614+   #ifndef STBI_ONLY_PNM
 615+   #define STBI_NO_PNM
 616+   #endif
 617+#endif
 618+
 619+#if defined(STBI_NO_PNG) && !defined(STBI_SUPPORT_ZLIB) && !defined(STBI_NO_ZLIB)
 620+#define STBI_NO_ZLIB
 621+#endif
 622+
 623+
 624+#include <stdarg.h>
 625+#include <stddef.h> /* ptrdiff_t on osx */
 626+#include <stdlib.h>
 627+#include <string.h>
 628+#include <limits.h>
 629+
 630+#if !defined(STBI_NO_LINEAR) || !defined(STBI_NO_HDR)
 631+#include <math.h> /* ldexp, pow */
 632+#endif
 633+
 634+#ifndef STBI_NO_STDIO
 635+#include <stdio.h>
 636+#endif
 637+
 638+#ifndef STBI_ASSERT
 639+#include <assert.h>
 640+#define STBI_ASSERT(x) assert(x)
 641+#endif
 642+
 643+#ifdef __cplusplus
 644+#define STBI_EXTERN extern "C"
 645+#else
 646+#define STBI_EXTERN extern
 647+#endif
 648+
 649+
 650+#ifndef _MSC_VER
 651+   #ifdef __cplusplus
 652+   #define stbi_inline inline
 653+   #else
 654+   #define stbi_inline
 655+   #endif
 656+#else
 657+   #define stbi_inline __forceinline
 658+#endif
 659+
 660+#ifndef STBI_NO_THREAD_LOCALS
 661+   #if defined(__cplusplus) &&  __cplusplus >= 201103L
 662+      #define STBI_THREAD_LOCAL       thread_local
 663+   #elif defined(__GNUC__) && __GNUC__ < 5
 664+      #define STBI_THREAD_LOCAL       __thread
 665+   #elif defined(_MSC_VER)
 666+      #define STBI_THREAD_LOCAL       __declspec(thread)
 667+   #elif defined (__STDC_VERSION__) && __STDC_VERSION__ >= 201112L && !defined(__STDC_NO_THREADS__)
 668+      #define STBI_THREAD_LOCAL       _Thread_local
 669+   #endif
 670+
 671+   #ifndef STBI_THREAD_LOCAL
 672+      #if defined(__GNUC__)
 673+        #define STBI_THREAD_LOCAL       __thread
 674+      #endif
 675+   #endif
 676+#endif
 677+
 678+#if defined(_MSC_VER) || defined(__SYMBIAN32__)
 679+typedef unsigned short stbi__uint16;
 680+typedef   signed short stbi__int16;
 681+typedef unsigned int   stbi__uint32;
 682+typedef   signed int   stbi__int32;
 683+#else
 684+#include <stdint.h>
 685+typedef uint16_t stbi__uint16;
 686+typedef int16_t  stbi__int16;
 687+typedef uint32_t stbi__uint32;
 688+typedef int32_t  stbi__int32;
 689+#endif
 690+
 691+/* should produce compiler error if size is wrong */
 692+typedef unsigned char validate_uint32[sizeof(stbi__uint32)==4 ? 1 : -1];
 693+
 694+#ifdef _MSC_VER
 695+#define STBI_NOTUSED(v)  (void)(v)
 696+#else
 697+#define STBI_NOTUSED(v)  (void)sizeof(v)
 698+#endif
 699+
 700+#ifdef _MSC_VER
 701+#define STBI_HAS_LROTL
 702+#endif
 703+
 704+#ifdef STBI_HAS_LROTL
 705+   #define stbi_lrot(x,y)  _lrotl(x,y)
 706+#else
 707+   #define stbi_lrot(x,y)  (((x) << (y)) | ((x) >> (-(y) & 31)))
 708+#endif
 709+
 710+#if defined(STBI_MALLOC) && defined(STBI_FREE) && (defined(STBI_REALLOC) || defined(STBI_REALLOC_SIZED))
 711+/* ok */
 712+#elif !defined(STBI_MALLOC) && !defined(STBI_FREE) && !defined(STBI_REALLOC) && !defined(STBI_REALLOC_SIZED)
 713+/* ok */
 714+#else
 715+#error "Must define all or none of STBI_MALLOC, STBI_FREE, and STBI_REALLOC (or STBI_REALLOC_SIZED)."
 716+#endif
 717+
 718+#ifndef STBI_MALLOC
 719+#define STBI_MALLOC(sz)           malloc(sz)
 720+#define STBI_REALLOC(p,newsz)     realloc(p,newsz)
 721+#define STBI_FREE(p)              free(p)
 722+#endif
 723+
 724+#ifndef STBI_REALLOC_SIZED
 725+#define STBI_REALLOC_SIZED(p,oldsz,newsz) STBI_REALLOC(p,newsz)
 726+#endif
 727+
 728+/* x86/x64 detection */
 729+#if defined(__x86_64__) || defined(_M_X64)
 730+#define STBI__X64_TARGET
 731+#elif defined(__i386) || defined(_M_IX86)
 732+#define STBI__X86_TARGET
 733+#endif
 734+
 735+#if defined(__GNUC__) && defined(STBI__X86_TARGET) && !defined(__SSE2__) && !defined(STBI_NO_SIMD)
 736+/*
 737+ * gcc doesn't support sse2 intrinsics unless you compile with -msse2,
 738+ * which in turn means it gets to use SSE2 everywhere. This is unfortunate,
 739+ * but previous attempts to provide the SSE2 functions with runtime
 740+ * detection caused numerous issues. The way architecture extensions are
 741+ * exposed in GCC/Clang is, sadly, not really suited for one-file libs.
 742+ * New behavior: if compiled with -msse2, we use SSE2 without any
 743+ * detection; if not, we don't use it at all.
 744+ */
 745+#define STBI_NO_SIMD
 746+#endif
 747+
 748+#if defined(__MINGW32__) && defined(STBI__X86_TARGET) && !defined(STBI_MINGW_ENABLE_SSE2) && !defined(STBI_NO_SIMD)
 749+/*
 750+ * Note that __MINGW32__ doesn't actually mean 32-bit, so we have to avoid STBI__X64_TARGET
 751+ *
 752+ * 32-bit MinGW wants ESP to be 16-byte aligned, but this is not in the
 753+ * Windows ABI and VC++ as well as Windows DLLs don't maintain that invariant.
 754+ * As a result, enabling SSE2 on 32-bit MinGW is dangerous when not
 755+ * simultaneously enabling "-mstackrealign".
 756+ *
 757+ * See https://github.com/nothings/stb/issues/81 for more information.
 758+ *
 759+ * So default to no SSE2 on 32-bit MinGW. If you've read this far and added
 760+ * -mstackrealign to your build settings, feel free to #define STBI_MINGW_ENABLE_SSE2.
 761+ */
 762+#define STBI_NO_SIMD
 763+#endif
 764+
 765+#if !defined(STBI_NO_SIMD) && (defined(STBI__X86_TARGET) || defined(STBI__X64_TARGET))
 766+#define STBI_SSE2
 767+#include <emmintrin.h>
 768+
 769+#ifdef _MSC_VER
 770+
 771+#if _MSC_VER >= 1400 /* not VC6 */
 772+#include <intrin.h> /* __cpuid */
 773+static int stbi__cpuid3(void)
 774+{
 775+   int info[4];
 776+   __cpuid(info,1);
 777+   return info[3];
 778+}
 779+#else
 780+static int stbi__cpuid3(void)
 781+{
 782+   int res;
 783+   __asm {
 784+      mov  eax,1
 785+      cpuid
 786+      mov  res,edx
 787+   }
 788+   return res;
 789+}
 790+#endif
 791+
 792+#define STBI_SIMD_ALIGN(type, name) __declspec(align(16)) type name
 793+
 794+#if !defined(STBI_NO_JPEG) && defined(STBI_SSE2)
 795+static int stbi__sse2_available(void)
 796+{
 797+   int info3 = stbi__cpuid3();
 798+   return ((info3 >> 26) & 1) != 0;
 799+}
 800+#endif
 801+
 802+#else /* assume GCC-style if not VC++ */
 803+#define STBI_SIMD_ALIGN(type, name) type name __attribute__((aligned(16)))
 804+
 805+#if !defined(STBI_NO_JPEG) && defined(STBI_SSE2)
 806+static int stbi__sse2_available(void)
 807+{
 808+   /*
 809+    * If we're even attempting to compile this on GCC/Clang, that means
 810+    * -msse2 is on, which means the compiler is allowed to use SSE2
 811+    * instructions at will, and so are we.
 812+    */
 813+   return 1;
 814+}
 815+#endif
 816+
 817+#endif
 818+#endif
 819+
 820+/* ARM NEON */
 821+#if defined(STBI_NO_SIMD) && defined(STBI_NEON)
 822+#undef STBI_NEON
 823+#endif
 824+
 825+#ifdef STBI_NEON
 826+#include <arm_neon.h>
 827+#ifdef _MSC_VER
 828+#define STBI_SIMD_ALIGN(type, name) __declspec(align(16)) type name
 829+#else
 830+#define STBI_SIMD_ALIGN(type, name) type name __attribute__((aligned(16)))
 831+#endif
 832+#endif
 833+
 834+#ifndef STBI_SIMD_ALIGN
 835+#define STBI_SIMD_ALIGN(type, name) type name
 836+#endif
 837+
 838+#ifndef STBI_MAX_DIMENSIONS
 839+#define STBI_MAX_DIMENSIONS (1 << 24)
 840+#endif
 841+
 842+/*/////////////////////////////////////////////*/
 843+/*
 844+ *
 845+ *  stbi__context struct and start_xxx functions
 846+ */
 847+
 848+/*
 849+ * stbi__context structure is our basic context used by all images, so it
 850+ * contains all the IO context, plus some basic image information
 851+ */
 852+typedef struct
 853+{
 854+   stbi__uint32 img_x, img_y;
 855+   int img_n, img_out_n;
 856+
 857+   stbi_io_callbacks io;
 858+   void *io_user_data;
 859+
 860+   int read_from_callbacks;
 861+   int buflen;
 862+   stbi_uc buffer_start[128];
 863+   int callback_already_read;
 864+
 865+   stbi_uc *img_buffer, *img_buffer_end;
 866+   stbi_uc *img_buffer_original, *img_buffer_original_end;
 867+} stbi__context;
 868+
 869+
 870+static void stbi__refill_buffer(stbi__context *s);
 871+
 872+/* initialize a memory-decode context */
 873+static void stbi__start_mem(stbi__context *s, stbi_uc const *buffer, int len)
 874+{
 875+   s->io.read = NULL;
 876+   s->read_from_callbacks = 0;
 877+   s->callback_already_read = 0;
 878+   s->img_buffer = s->img_buffer_original = (stbi_uc *) buffer;
 879+   s->img_buffer_end = s->img_buffer_original_end = (stbi_uc *) buffer+len;
 880+}
 881+
 882+/* initialize a callback-based context */
 883+static void stbi__start_callbacks(stbi__context *s, stbi_io_callbacks *c, void *user)
 884+{
 885+   s->io = *c;
 886+   s->io_user_data = user;
 887+   s->buflen = sizeof(s->buffer_start);
 888+   s->read_from_callbacks = 1;
 889+   s->callback_already_read = 0;
 890+   s->img_buffer = s->img_buffer_original = s->buffer_start;
 891+   stbi__refill_buffer(s);
 892+   s->img_buffer_original_end = s->img_buffer_end;
 893+}
 894+
 895+#ifndef STBI_NO_STDIO
 896+
 897+static int stbi__stdio_read(void *user, char *data, int size)
 898+{
 899+   return (int) fread(data,1,size,(FILE*) user);
 900+}
 901+
 902+static void stbi__stdio_skip(void *user, int n)
 903+{
 904+   int ch;
 905+   fseek((FILE*) user, n, SEEK_CUR);
 906+   ch = fgetc((FILE*) user);  /* have to read a byte to reset feof()'s flag */
 907+   if (ch != EOF) {
 908+      ungetc(ch, (FILE *) user);  /* push byte back onto stream if valid. */
 909+   }
 910+}
 911+
 912+static int stbi__stdio_eof(void *user)
 913+{
 914+   return feof((FILE*) user) || ferror((FILE *) user);
 915+}
 916+
 917+static stbi_io_callbacks stbi__stdio_callbacks =
 918+{
 919+   stbi__stdio_read,
 920+   stbi__stdio_skip,
 921+   stbi__stdio_eof,
 922+};
 923+
 924+static void stbi__start_file(stbi__context *s, FILE *f)
 925+{
 926+   stbi__start_callbacks(s, &stbi__stdio_callbacks, (void *) f);
 927+}
 928+
 929+/* static void stop_file(stbi__context *s) { } */
 930+
 931+#endif /* !STBI_NO_STDIO */
 932+
 933+static void stbi__rewind(stbi__context *s)
 934+{
 935+   /*
 936+    * conceptually rewind SHOULD rewind to the beginning of the stream,
 937+    * but we just rewind to the beginning of the initial buffer, because
 938+    * we only use it after doing 'test', which only ever looks at at most 92 bytes
 939+    */
 940+   s->img_buffer = s->img_buffer_original;
 941+   s->img_buffer_end = s->img_buffer_original_end;
 942+}
 943+
 944+enum
 945+{
 946+   STBI_ORDER_RGB,
 947+   STBI_ORDER_BGR
 948+};
 949+
 950+typedef struct
 951+{
 952+   int bits_per_channel;
 953+   int num_channels;
 954+   int channel_order;
 955+} stbi__result_info;
 956+
 957+#ifndef STBI_NO_JPEG
 958+static int      stbi__jpeg_test(stbi__context *s);
 959+static void    *stbi__jpeg_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
 960+static int      stbi__jpeg_info(stbi__context *s, int *x, int *y, int *comp);
 961+#endif
 962+
 963+#ifndef STBI_NO_PNG
 964+static int      stbi__png_test(stbi__context *s);
 965+static void    *stbi__png_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
 966+static int      stbi__png_info(stbi__context *s, int *x, int *y, int *comp);
 967+static int      stbi__png_is16(stbi__context *s);
 968+#endif
 969+
 970+#ifndef STBI_NO_BMP
 971+static int      stbi__bmp_test(stbi__context *s);
 972+static void    *stbi__bmp_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
 973+static int      stbi__bmp_info(stbi__context *s, int *x, int *y, int *comp);
 974+#endif
 975+
 976+#ifndef STBI_NO_TGA
 977+static int      stbi__tga_test(stbi__context *s);
 978+static void    *stbi__tga_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
 979+static int      stbi__tga_info(stbi__context *s, int *x, int *y, int *comp);
 980+#endif
 981+
 982+#ifndef STBI_NO_PSD
 983+static int      stbi__psd_test(stbi__context *s);
 984+static void    *stbi__psd_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri, int bpc);
 985+static int      stbi__psd_info(stbi__context *s, int *x, int *y, int *comp);
 986+static int      stbi__psd_is16(stbi__context *s);
 987+#endif
 988+
 989+#ifndef STBI_NO_HDR
 990+static int      stbi__hdr_test(stbi__context *s);
 991+static float   *stbi__hdr_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
 992+static int      stbi__hdr_info(stbi__context *s, int *x, int *y, int *comp);
 993+#endif
 994+
 995+#ifndef STBI_NO_PIC
 996+static int      stbi__pic_test(stbi__context *s);
 997+static void    *stbi__pic_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
 998+static int      stbi__pic_info(stbi__context *s, int *x, int *y, int *comp);
 999+#endif
1000+
1001+#ifndef STBI_NO_GIF
1002+static int      stbi__gif_test(stbi__context *s);
1003+static void    *stbi__gif_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
1004+static void    *stbi__load_gif_main(stbi__context *s, int **delays, int *x, int *y, int *z, int *comp, int req_comp);
1005+static int      stbi__gif_info(stbi__context *s, int *x, int *y, int *comp);
1006+#endif
1007+
1008+#ifndef STBI_NO_PNM
1009+static int      stbi__pnm_test(stbi__context *s);
1010+static void    *stbi__pnm_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri);
1011+static int      stbi__pnm_info(stbi__context *s, int *x, int *y, int *comp);
1012+static int      stbi__pnm_is16(stbi__context *s);
1013+#endif
1014+
1015+static
1016+#ifdef STBI_THREAD_LOCAL
1017+STBI_THREAD_LOCAL
1018+#endif
1019+const char *stbi__g_failure_reason;
1020+
1021+STBIDEF const char *stbi_failure_reason(void)
1022+{
1023+   return stbi__g_failure_reason;
1024+}
1025+
1026+#ifndef STBI_NO_FAILURE_STRINGS
1027+static int stbi__err(const char *str)
1028+{
1029+   stbi__g_failure_reason = str;
1030+   return 0;
1031+}
1032+#endif
1033+
1034+static void *stbi__malloc(size_t size)
1035+{
1036+    return STBI_MALLOC(size);
1037+}
1038+
1039+/*
1040+ * stb_image uses ints pervasively, including for offset calculations.
1041+ * therefore the largest decoded image size we can support with the
1042+ * current code, even on 64-bit targets, is INT_MAX. this is not a
1043+ * significant limitation for the intended use case.
1044+ *
1045+ * we do, however, need to make sure our size calculations don't
1046+ * overflow. hence a few helper functions for size calculations that
1047+ * multiply integers together, making sure that they're non-negative
1048+ * and no overflow occurs.
1049+ */
1050+
1051+/*
1052+ * return 1 if the sum is valid, 0 on overflow.
1053+ * negative terms are considered invalid.
1054+ */
1055+static int stbi__addsizes_valid(int a, int b)
1056+{
1057+   if (b < 0) return 0;
1058+   /*
1059+    * now 0 <= b <= INT_MAX, hence also
1060+    * 0 <= INT_MAX - b <= INTMAX.
1061+    * And "a + b <= INT_MAX" (which might overflow) is the
1062+    * same as a <= INT_MAX - b (no overflow)
1063+    */
1064+   return a <= INT_MAX - b;
1065+}
1066+
1067+/*
1068+ * returns 1 if the product is valid, 0 on overflow.
1069+ * negative factors are considered invalid.
1070+ */
1071+static int stbi__mul2sizes_valid(int a, int b)
1072+{
1073+   if (a < 0 || b < 0) return 0;
1074+   if (b == 0) return 1; /* mul-by-0 is always safe */
1075+   /* portable way to check for no overflows in a*b */
1076+   return a <= INT_MAX/b;
1077+}
1078+
1079+#if !defined(STBI_NO_JPEG) || !defined(STBI_NO_PNG) || !defined(STBI_NO_TGA) || !defined(STBI_NO_HDR)
1080+/* returns 1 if "a*b + add" has no negative terms/factors and doesn't overflow */
1081+static int stbi__mad2sizes_valid(int a, int b, int add)
1082+{
1083+   return stbi__mul2sizes_valid(a, b) && stbi__addsizes_valid(a*b, add);
1084+}
1085+#endif
1086+
1087+/* returns 1 if "a*b*c + add" has no negative terms/factors and doesn't overflow */
1088+static int stbi__mad3sizes_valid(int a, int b, int c, int add)
1089+{
1090+   return stbi__mul2sizes_valid(a, b) && stbi__mul2sizes_valid(a*b, c) &&
1091+      stbi__addsizes_valid(a*b*c, add);
1092+}
1093+
1094+/* returns 1 if "a*b*c*d + add" has no negative terms/factors and doesn't overflow */
1095+#if !defined(STBI_NO_LINEAR) || !defined(STBI_NO_HDR) || !defined(STBI_NO_PNM)
1096+static int stbi__mad4sizes_valid(int a, int b, int c, int d, int add)
1097+{
1098+   return stbi__mul2sizes_valid(a, b) && stbi__mul2sizes_valid(a*b, c) &&
1099+      stbi__mul2sizes_valid(a*b*c, d) && stbi__addsizes_valid(a*b*c*d, add);
1100+}
1101+#endif
1102+
1103+#if !defined(STBI_NO_JPEG) || !defined(STBI_NO_PNG) || !defined(STBI_NO_TGA) || !defined(STBI_NO_HDR)
1104+/* mallocs with size overflow checking */
1105+static void *stbi__malloc_mad2(int a, int b, int add)
1106+{
1107+   if (!stbi__mad2sizes_valid(a, b, add)) return NULL;
1108+   return stbi__malloc(a*b + add);
1109+}
1110+#endif
1111+
1112+static void *stbi__malloc_mad3(int a, int b, int c, int add)
1113+{
1114+   if (!stbi__mad3sizes_valid(a, b, c, add)) return NULL;
1115+   return stbi__malloc(a*b*c + add);
1116+}
1117+
1118+#if !defined(STBI_NO_LINEAR) || !defined(STBI_NO_HDR) || !defined(STBI_NO_PNM)
1119+static void *stbi__malloc_mad4(int a, int b, int c, int d, int add)
1120+{
1121+   if (!stbi__mad4sizes_valid(a, b, c, d, add)) return NULL;
1122+   return stbi__malloc(a*b*c*d + add);
1123+}
1124+#endif
1125+
1126+/* returns 1 if the sum of two signed ints is valid (between -2^31 and 2^31-1 inclusive), 0 on overflow. */
1127+static int stbi__addints_valid(int a, int b)
1128+{
1129+   if ((a >= 0) != (b >= 0)) return 1; /* a and b have different signs, so no overflow */
1130+   if (a < 0 && b < 0) return a >= INT_MIN - b; /* same as a + b >= INT_MIN; INT_MIN - b cannot overflow since b < 0. */
1131+   return a <= INT_MAX - b;
1132+}
1133+
1134+/* returns 1 if the product of two ints fits in a signed short, 0 on overflow. */
1135+static int stbi__mul2shorts_valid(int a, int b)
1136+{
1137+   if (b == 0 || b == -1) return 1; /* multiplication by 0 is always 0; check for -1 so SHRT_MIN/b doesn't overflow */
1138+   if ((a >= 0) == (b >= 0)) return a <= SHRT_MAX/b; /* product is positive, so similar to mul2sizes_valid */
1139+   if (b < 0) return a <= SHRT_MIN / b; /* same as a * b >= SHRT_MIN */
1140+   return a >= SHRT_MIN / b;
1141+}
1142+
1143+/*
1144+ * stbi__err - error
1145+ * stbi__errpf - error returning pointer to float
1146+ * stbi__errpuc - error returning pointer to unsigned char
1147+ */
1148+
1149+#ifdef STBI_NO_FAILURE_STRINGS
1150+   #define stbi__err(x,y)  0
1151+#elif defined(STBI_FAILURE_USERMSG)
1152+   #define stbi__err(x,y)  stbi__err(y)
1153+#else
1154+   #define stbi__err(x,y)  stbi__err(x)
1155+#endif
1156+
1157+#define stbi__errpf(x,y)   ((float *)(size_t) (stbi__err(x,y)?NULL:NULL))
1158+#define stbi__errpuc(x,y)  ((unsigned char *)(size_t) (stbi__err(x,y)?NULL:NULL))
1159+
1160+STBIDEF void stbi_image_free(void *retval_from_stbi_load)
1161+{
1162+   STBI_FREE(retval_from_stbi_load);
1163+}
1164+
1165+#ifndef STBI_NO_LINEAR
1166+static float   *stbi__ldr_to_hdr(stbi_uc *data, int x, int y, int comp);
1167+#endif
1168+
1169+#ifndef STBI_NO_HDR
1170+static stbi_uc *stbi__hdr_to_ldr(float   *data, int x, int y, int comp);
1171+#endif
1172+
1173+static int stbi__vertically_flip_on_load_global = 0;
1174+
1175+STBIDEF void stbi_set_flip_vertically_on_load(int flag_true_if_should_flip)
1176+{
1177+   stbi__vertically_flip_on_load_global = flag_true_if_should_flip;
1178+}
1179+
1180+#ifndef STBI_THREAD_LOCAL
1181+#define stbi__vertically_flip_on_load  stbi__vertically_flip_on_load_global
1182+#else
1183+static STBI_THREAD_LOCAL int stbi__vertically_flip_on_load_local, stbi__vertically_flip_on_load_set;
1184+
1185+STBIDEF void stbi_set_flip_vertically_on_load_thread(int flag_true_if_should_flip)
1186+{
1187+   stbi__vertically_flip_on_load_local = flag_true_if_should_flip;
1188+   stbi__vertically_flip_on_load_set = 1;
1189+}
1190+
1191+#define stbi__vertically_flip_on_load  (stbi__vertically_flip_on_load_set       \
1192+                                         ? stbi__vertically_flip_on_load_local  \
1193+                                         : stbi__vertically_flip_on_load_global)
1194+#endif /* STBI_THREAD_LOCAL */
1195+
1196+static void *stbi__load_main(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri, int bpc)
1197+{
1198+   memset(ri, 0, sizeof(*ri)); /* make sure it's initialized if we add new fields */
1199+   ri->bits_per_channel = 8; /* default is 8 so most paths don't have to be changed */
1200+   ri->channel_order = STBI_ORDER_RGB; /* all current input & output are this, but this is here so we can add BGR order */
1201+   ri->num_channels = 0;
1202+
1203+   /*
1204+    * test the formats with a very explicit header first (at least a FOURCC
1205+    * or distinctive magic number first)
1206+    */
1207+   #ifndef STBI_NO_PNG
1208+   if (stbi__png_test(s))  return stbi__png_load(s,x,y,comp,req_comp, ri);
1209+   #endif
1210+   #ifndef STBI_NO_BMP
1211+   if (stbi__bmp_test(s))  return stbi__bmp_load(s,x,y,comp,req_comp, ri);
1212+   #endif
1213+   #ifndef STBI_NO_GIF
1214+   if (stbi__gif_test(s))  return stbi__gif_load(s,x,y,comp,req_comp, ri);
1215+   #endif
1216+   #ifndef STBI_NO_PSD
1217+   if (stbi__psd_test(s))  return stbi__psd_load(s,x,y,comp,req_comp, ri, bpc);
1218+   #else
1219+   STBI_NOTUSED(bpc);
1220+   #endif
1221+   #ifndef STBI_NO_PIC
1222+   if (stbi__pic_test(s))  return stbi__pic_load(s,x,y,comp,req_comp, ri);
1223+   #endif
1224+
1225+   /*
1226+    * then the formats that can end up attempting to load with just 1 or 2
1227+    * bytes matching expectations; these are prone to false positives, so
1228+    * try them later
1229+    */
1230+   #ifndef STBI_NO_JPEG
1231+   if (stbi__jpeg_test(s)) return stbi__jpeg_load(s,x,y,comp,req_comp, ri);
1232+   #endif
1233+   #ifndef STBI_NO_PNM
1234+   if (stbi__pnm_test(s))  return stbi__pnm_load(s,x,y,comp,req_comp, ri);
1235+   #endif
1236+
1237+   #ifndef STBI_NO_HDR
1238+   if (stbi__hdr_test(s)) {
1239+      float *hdr = stbi__hdr_load(s, x,y,comp,req_comp, ri);
1240+      return stbi__hdr_to_ldr(hdr, *x, *y, req_comp ? req_comp : *comp);
1241+   }
1242+   #endif
1243+
1244+   #ifndef STBI_NO_TGA
1245+   /* test tga last because it's a crappy test! */
1246+   if (stbi__tga_test(s))
1247+      return stbi__tga_load(s,x,y,comp,req_comp, ri);
1248+   #endif
1249+
1250+   return stbi__errpuc("unknown image type", "Image not of any known type, or corrupt");
1251+}
1252+
1253+static stbi_uc *stbi__convert_16_to_8(stbi__uint16 *orig, int w, int h, int channels)
1254+{
1255+   int i;
1256+   int img_len = w * h * channels;
1257+   stbi_uc *reduced;
1258+
1259+   reduced = (stbi_uc *) stbi__malloc(img_len);
1260+   if (reduced == NULL) return stbi__errpuc("outofmem", "Out of memory");
1261+
1262+   for (i = 0; i < img_len; ++i)
1263+      reduced[i] = (stbi_uc)((orig[i] >> 8) & 0xFF); /* top half of each byte is sufficient approx of 16->8 bit scaling */
1264+
1265+   STBI_FREE(orig);
1266+   return reduced;
1267+}
1268+
1269+static stbi__uint16 *stbi__convert_8_to_16(stbi_uc *orig, int w, int h, int channels)
1270+{
1271+   int i;
1272+   int img_len = w * h * channels;
1273+   stbi__uint16 *enlarged;
1274+
1275+   enlarged = (stbi__uint16 *) stbi__malloc(img_len*2);
1276+   if (enlarged == NULL) return (stbi__uint16 *) stbi__errpuc("outofmem", "Out of memory");
1277+
1278+   for (i = 0; i < img_len; ++i)
1279+      enlarged[i] = (stbi__uint16)((orig[i] << 8) + orig[i]); /* replicate to high and low byte, maps 0->0, 255->0xffff */
1280+
1281+   STBI_FREE(orig);
1282+   return enlarged;
1283+}
1284+
1285+static void stbi__vertical_flip(void *image, int w, int h, int bytes_per_pixel)
1286+{
1287+   int row;
1288+   size_t bytes_per_row = (size_t)w * bytes_per_pixel;
1289+   stbi_uc temp[2048];
1290+   stbi_uc *bytes = (stbi_uc *)image;
1291+
1292+   for (row = 0; row < (h>>1); row++) {
1293+      stbi_uc *row0 = bytes + row*bytes_per_row;
1294+      stbi_uc *row1 = bytes + (h - row - 1)*bytes_per_row;
1295+      /* swap row0 with row1 */
1296+      size_t bytes_left = bytes_per_row;
1297+      while (bytes_left) {
1298+         size_t bytes_copy = (bytes_left < sizeof(temp)) ? bytes_left : sizeof(temp);
1299+         memcpy(temp, row0, bytes_copy);
1300+         memcpy(row0, row1, bytes_copy);
1301+         memcpy(row1, temp, bytes_copy);
1302+         row0 += bytes_copy;
1303+         row1 += bytes_copy;
1304+         bytes_left -= bytes_copy;
1305+      }
1306+   }
1307+}
1308+
1309+#ifndef STBI_NO_GIF
1310+static void stbi__vertical_flip_slices(void *image, int w, int h, int z, int bytes_per_pixel)
1311+{
1312+   int slice;
1313+   int slice_size = w * h * bytes_per_pixel;
1314+
1315+   stbi_uc *bytes = (stbi_uc *)image;
1316+   for (slice = 0; slice < z; ++slice) {
1317+      stbi__vertical_flip(bytes, w, h, bytes_per_pixel);
1318+      bytes += slice_size;
1319+   }
1320+}
1321+#endif
1322+
1323+static unsigned char *stbi__load_and_postprocess_8bit(stbi__context *s, int *x, int *y, int *comp, int req_comp)
1324+{
1325+   stbi__result_info ri;
1326+   void *result = stbi__load_main(s, x, y, comp, req_comp, &ri, 8);
1327+
1328+   if (result == NULL)
1329+      return NULL;
1330+
1331+   /* it is the responsibility of the loaders to make sure we get either 8 or 16 bit. */
1332+   STBI_ASSERT(ri.bits_per_channel == 8 || ri.bits_per_channel == 16);
1333+
1334+   if (ri.bits_per_channel != 8) {
1335+      result = stbi__convert_16_to_8((stbi__uint16 *) result, *x, *y, req_comp == 0 ? *comp : req_comp);
1336+      ri.bits_per_channel = 8;
1337+   }
1338+
1339+   /* @TODO: move stbi__convert_format to here */
1340+
1341+   if (stbi__vertically_flip_on_load) {
1342+      int channels = req_comp ? req_comp : *comp;
1343+      stbi__vertical_flip(result, *x, *y, channels * sizeof(stbi_uc));
1344+   }
1345+
1346+   return (unsigned char *) result;
1347+}
1348+
1349+static stbi__uint16 *stbi__load_and_postprocess_16bit(stbi__context *s, int *x, int *y, int *comp, int req_comp)
1350+{
1351+   stbi__result_info ri;
1352+   void *result = stbi__load_main(s, x, y, comp, req_comp, &ri, 16);
1353+
1354+   if (result == NULL)
1355+      return NULL;
1356+
1357+   /* it is the responsibility of the loaders to make sure we get either 8 or 16 bit. */
1358+   STBI_ASSERT(ri.bits_per_channel == 8 || ri.bits_per_channel == 16);
1359+
1360+   if (ri.bits_per_channel != 16) {
1361+      result = stbi__convert_8_to_16((stbi_uc *) result, *x, *y, req_comp == 0 ? *comp : req_comp);
1362+      ri.bits_per_channel = 16;
1363+   }
1364+
1365+   /*
1366+    * @TODO: move stbi__convert_format16 to here
1367+    * @TODO: special case RGB-to-Y (and RGBA-to-YA) for 8-bit-to-16-bit case to keep more precision
1368+    */
1369+
1370+   if (stbi__vertically_flip_on_load) {
1371+      int channels = req_comp ? req_comp : *comp;
1372+      stbi__vertical_flip(result, *x, *y, channels * sizeof(stbi__uint16));
1373+   }
1374+
1375+   return (stbi__uint16 *) result;
1376+}
1377+
1378+#if !defined(STBI_NO_HDR) && !defined(STBI_NO_LINEAR)
1379+static void stbi__float_postprocess(float *result, int *x, int *y, int *comp, int req_comp)
1380+{
1381+   if (stbi__vertically_flip_on_load && result != NULL) {
1382+      int channels = req_comp ? req_comp : *comp;
1383+      stbi__vertical_flip(result, *x, *y, channels * sizeof(float));
1384+   }
1385+}
1386+#endif
1387+
1388+#ifndef STBI_NO_STDIO
1389+
1390+#if defined(_WIN32) && defined(STBI_WINDOWS_UTF8)
1391+STBI_EXTERN __declspec(dllimport) int __stdcall MultiByteToWideChar(unsigned int cp, unsigned long flags, const char *str, int cbmb, wchar_t *widestr, int cchwide);
1392+STBI_EXTERN __declspec(dllimport) int __stdcall WideCharToMultiByte(unsigned int cp, unsigned long flags, const wchar_t *widestr, int cchwide, char *str, int cbmb, const char *defchar, int *used_default);
1393+#endif
1394+
1395+#if defined(_WIN32) && defined(STBI_WINDOWS_UTF8)
1396+STBIDEF int stbi_convert_wchar_to_utf8(char *buffer, size_t bufferlen, const wchar_t* input)
1397+{
1398+	return WideCharToMultiByte(65001 /* UTF8 */, 0, input, -1, buffer, (int) bufferlen, NULL, NULL);
1399+}
1400+#endif
1401+
1402+static FILE *stbi__fopen(char const *filename, char const *mode)
1403+{
1404+   FILE *f;
1405+#if defined(_WIN32) && defined(STBI_WINDOWS_UTF8)
1406+   wchar_t wMode[64];
1407+   wchar_t wFilename[1024];
1408+	if (0 == MultiByteToWideChar(65001 /* UTF8 */, 0, filename, -1, wFilename, sizeof(wFilename)/sizeof(*wFilename)))
1409+      return 0;
1410+
1411+	if (0 == MultiByteToWideChar(65001 /* UTF8 */, 0, mode, -1, wMode, sizeof(wMode)/sizeof(*wMode)))
1412+      return 0;
1413+
1414+#if defined(_MSC_VER) && _MSC_VER >= 1400
1415+	if (0 != _wfopen_s(&f, wFilename, wMode))
1416+		f = 0;
1417+#else
1418+   f = _wfopen(wFilename, wMode);
1419+#endif
1420+
1421+#elif defined(_MSC_VER) && _MSC_VER >= 1400
1422+   if (0 != fopen_s(&f, filename, mode))
1423+      f=0;
1424+#else
1425+   f = fopen(filename, mode);
1426+#endif
1427+   return f;
1428+}
1429+
1430+
1431+STBIDEF stbi_uc *stbi_load(char const *filename, int *x, int *y, int *comp, int req_comp)
1432+{
1433+   FILE *f = stbi__fopen(filename, "rb");
1434+   unsigned char *result;
1435+   if (!f) return stbi__errpuc("can't fopen", "Unable to open file");
1436+   result = stbi_load_from_file(f,x,y,comp,req_comp);
1437+   fclose(f);
1438+   return result;
1439+}
1440+
1441+STBIDEF stbi_uc *stbi_load_from_file(FILE *f, int *x, int *y, int *comp, int req_comp)
1442+{
1443+   unsigned char *result;
1444+   stbi__context s;
1445+   stbi__start_file(&s,f);
1446+   result = stbi__load_and_postprocess_8bit(&s,x,y,comp,req_comp);
1447+   if (result) {
1448+      /* need to 'unget' all the characters in the IO buffer */
1449+      fseek(f, - (int) (s.img_buffer_end - s.img_buffer), SEEK_CUR);
1450+   }
1451+   return result;
1452+}
1453+
1454+STBIDEF stbi__uint16 *stbi_load_from_file_16(FILE *f, int *x, int *y, int *comp, int req_comp)
1455+{
1456+   stbi__uint16 *result;
1457+   stbi__context s;
1458+   stbi__start_file(&s,f);
1459+   result = stbi__load_and_postprocess_16bit(&s,x,y,comp,req_comp);
1460+   if (result) {
1461+      /* need to 'unget' all the characters in the IO buffer */
1462+      fseek(f, - (int) (s.img_buffer_end - s.img_buffer), SEEK_CUR);
1463+   }
1464+   return result;
1465+}
1466+
1467+STBIDEF stbi_us *stbi_load_16(char const *filename, int *x, int *y, int *comp, int req_comp)
1468+{
1469+   FILE *f = stbi__fopen(filename, "rb");
1470+   stbi__uint16 *result;
1471+   if (!f) return (stbi_us *) stbi__errpuc("can't fopen", "Unable to open file");
1472+   result = stbi_load_from_file_16(f,x,y,comp,req_comp);
1473+   fclose(f);
1474+   return result;
1475+}
1476+
1477+
1478+#endif /* !STBI_NO_STDIO */
1479+
1480+STBIDEF stbi_us *stbi_load_16_from_memory(stbi_uc const *buffer, int len, int *x, int *y, int *channels_in_file, int desired_channels)
1481+{
1482+   stbi__context s;
1483+   stbi__start_mem(&s,buffer,len);
1484+   return stbi__load_and_postprocess_16bit(&s,x,y,channels_in_file,desired_channels);
1485+}
1486+
1487+STBIDEF stbi_us *stbi_load_16_from_callbacks(stbi_io_callbacks const *clbk, void *user, int *x, int *y, int *channels_in_file, int desired_channels)
1488+{
1489+   stbi__context s;
1490+   stbi__start_callbacks(&s, (stbi_io_callbacks *)clbk, user);
1491+   return stbi__load_and_postprocess_16bit(&s,x,y,channels_in_file,desired_channels);
1492+}
1493+
1494+STBIDEF stbi_uc *stbi_load_from_memory(stbi_uc const *buffer, int len, int *x, int *y, int *comp, int req_comp)
1495+{
1496+   stbi__context s;
1497+   stbi__start_mem(&s,buffer,len);
1498+   return stbi__load_and_postprocess_8bit(&s,x,y,comp,req_comp);
1499+}
1500+
1501+STBIDEF stbi_uc *stbi_load_from_callbacks(stbi_io_callbacks const *clbk, void *user, int *x, int *y, int *comp, int req_comp)
1502+{
1503+   stbi__context s;
1504+   stbi__start_callbacks(&s, (stbi_io_callbacks *) clbk, user);
1505+   return stbi__load_and_postprocess_8bit(&s,x,y,comp,req_comp);
1506+}
1507+
1508+#ifndef STBI_NO_GIF
1509+STBIDEF stbi_uc *stbi_load_gif_from_memory(stbi_uc const *buffer, int len, int **delays, int *x, int *y, int *z, int *comp, int req_comp)
1510+{
1511+   unsigned char *result;
1512+   stbi__context s;
1513+   stbi__start_mem(&s,buffer,len);
1514+
1515+   result = (unsigned char*) stbi__load_gif_main(&s, delays, x, y, z, comp, req_comp);
1516+   if (stbi__vertically_flip_on_load) {
1517+      stbi__vertical_flip_slices( result, *x, *y, *z, *comp );
1518+   }
1519+
1520+   return result;
1521+}
1522+#endif
1523+
1524+#ifndef STBI_NO_LINEAR
1525+static float *stbi__loadf_main(stbi__context *s, int *x, int *y, int *comp, int req_comp)
1526+{
1527+   unsigned char *data;
1528+   #ifndef STBI_NO_HDR
1529+   if (stbi__hdr_test(s)) {
1530+      stbi__result_info ri;
1531+      float *hdr_data = stbi__hdr_load(s,x,y,comp,req_comp, &ri);
1532+      if (hdr_data)
1533+         stbi__float_postprocess(hdr_data,x,y,comp,req_comp);
1534+      return hdr_data;
1535+   }
1536+   #endif
1537+   data = stbi__load_and_postprocess_8bit(s, x, y, comp, req_comp);
1538+   if (data)
1539+      return stbi__ldr_to_hdr(data, *x, *y, req_comp ? req_comp : *comp);
1540+   return stbi__errpf("unknown image type", "Image not of any known type, or corrupt");
1541+}
1542+
1543+STBIDEF float *stbi_loadf_from_memory(stbi_uc const *buffer, int len, int *x, int *y, int *comp, int req_comp)
1544+{
1545+   stbi__context s;
1546+   stbi__start_mem(&s,buffer,len);
1547+   return stbi__loadf_main(&s,x,y,comp,req_comp);
1548+}
1549+
1550+STBIDEF float *stbi_loadf_from_callbacks(stbi_io_callbacks const *clbk, void *user, int *x, int *y, int *comp, int req_comp)
1551+{
1552+   stbi__context s;
1553+   stbi__start_callbacks(&s, (stbi_io_callbacks *) clbk, user);
1554+   return stbi__loadf_main(&s,x,y,comp,req_comp);
1555+}
1556+
1557+#ifndef STBI_NO_STDIO
1558+STBIDEF float *stbi_loadf(char const *filename, int *x, int *y, int *comp, int req_comp)
1559+{
1560+   float *result;
1561+   FILE *f = stbi__fopen(filename, "rb");
1562+   if (!f) return stbi__errpf("can't fopen", "Unable to open file");
1563+   result = stbi_loadf_from_file(f,x,y,comp,req_comp);
1564+   fclose(f);
1565+   return result;
1566+}
1567+
1568+STBIDEF float *stbi_loadf_from_file(FILE *f, int *x, int *y, int *comp, int req_comp)
1569+{
1570+   stbi__context s;
1571+   stbi__start_file(&s,f);
1572+   return stbi__loadf_main(&s,x,y,comp,req_comp);
1573+}
1574+#endif /* !STBI_NO_STDIO */
1575+
1576+#endif /* !STBI_NO_LINEAR */
1577+
1578+/*
1579+ * these is-hdr-or-not is defined independent of whether STBI_NO_LINEAR is
1580+ * defined, for API simplicity; if STBI_NO_LINEAR is defined, it always
1581+ * reports false!
1582+ */
1583+
1584+STBIDEF int stbi_is_hdr_from_memory(stbi_uc const *buffer, int len)
1585+{
1586+   #ifndef STBI_NO_HDR
1587+   stbi__context s;
1588+   stbi__start_mem(&s,buffer,len);
1589+   return stbi__hdr_test(&s);
1590+   #else
1591+   STBI_NOTUSED(buffer);
1592+   STBI_NOTUSED(len);
1593+   return 0;
1594+   #endif
1595+}
1596+
1597+#ifndef STBI_NO_STDIO
1598+STBIDEF int      stbi_is_hdr          (char const *filename)
1599+{
1600+   FILE *f = stbi__fopen(filename, "rb");
1601+   int result=0;
1602+   if (f) {
1603+      result = stbi_is_hdr_from_file(f);
1604+      fclose(f);
1605+   }
1606+   return result;
1607+}
1608+
1609+STBIDEF int stbi_is_hdr_from_file(FILE *f)
1610+{
1611+   #ifndef STBI_NO_HDR
1612+   long pos = ftell(f);
1613+   int res;
1614+   stbi__context s;
1615+   stbi__start_file(&s,f);
1616+   res = stbi__hdr_test(&s);
1617+   fseek(f, pos, SEEK_SET);
1618+   return res;
1619+   #else
1620+   STBI_NOTUSED(f);
1621+   return 0;
1622+   #endif
1623+}
1624+#endif /* !STBI_NO_STDIO */
1625+
1626+STBIDEF int      stbi_is_hdr_from_callbacks(stbi_io_callbacks const *clbk, void *user)
1627+{
1628+   #ifndef STBI_NO_HDR
1629+   stbi__context s;
1630+   stbi__start_callbacks(&s, (stbi_io_callbacks *) clbk, user);
1631+   return stbi__hdr_test(&s);
1632+   #else
1633+   STBI_NOTUSED(clbk);
1634+   STBI_NOTUSED(user);
1635+   return 0;
1636+   #endif
1637+}
1638+
1639+#ifndef STBI_NO_LINEAR
1640+static float stbi__l2h_gamma=2.2f, stbi__l2h_scale=1.0f;
1641+
1642+STBIDEF void   stbi_ldr_to_hdr_gamma(float gamma) { stbi__l2h_gamma = gamma; }
1643+STBIDEF void   stbi_ldr_to_hdr_scale(float scale) { stbi__l2h_scale = scale; }
1644+#endif
1645+
1646+static float stbi__h2l_gamma_i=1.0f/2.2f, stbi__h2l_scale_i=1.0f;
1647+
1648+STBIDEF void   stbi_hdr_to_ldr_gamma(float gamma) { stbi__h2l_gamma_i = 1/gamma; }
1649+STBIDEF void   stbi_hdr_to_ldr_scale(float scale) { stbi__h2l_scale_i = 1/scale; }
1650+
1651+
1652+/*////////////////////////////////////////////////////////////////////////////*/
1653+/*
1654+ *
1655+ * Common code used by all image loaders
1656+ *
1657+ */
1658+
1659+enum
1660+{
1661+   STBI__SCAN_load=0,
1662+   STBI__SCAN_type,
1663+   STBI__SCAN_header
1664+};
1665+
1666+static void stbi__refill_buffer(stbi__context *s)
1667+{
1668+   int n = (s->io.read)(s->io_user_data,(char*)s->buffer_start,s->buflen);
1669+   s->callback_already_read += (int) (s->img_buffer - s->img_buffer_original);
1670+   if (n == 0) {
1671+      /*
1672+       * at end of file, treat same as if from memory, but need to handle case
1673+       * where s->img_buffer isn't pointing to safe memory, e.g. 0-byte file
1674+       */
1675+      s->read_from_callbacks = 0;
1676+      s->img_buffer = s->buffer_start;
1677+      s->img_buffer_end = s->buffer_start+1;
1678+      *s->img_buffer = 0;
1679+   } else {
1680+      s->img_buffer = s->buffer_start;
1681+      s->img_buffer_end = s->buffer_start + n;
1682+   }
1683+}
1684+
1685+stbi_inline static stbi_uc stbi__get8(stbi__context *s)
1686+{
1687+   if (s->img_buffer < s->img_buffer_end)
1688+      return *s->img_buffer++;
1689+   if (s->read_from_callbacks) {
1690+      stbi__refill_buffer(s);
1691+      return *s->img_buffer++;
1692+   }
1693+   return 0;
1694+}
1695+
1696+#if defined(STBI_NO_JPEG) && defined(STBI_NO_HDR) && defined(STBI_NO_PIC) && defined(STBI_NO_PNM)
1697+/* nothing */
1698+#else
1699+stbi_inline static int stbi__at_eof(stbi__context *s)
1700+{
1701+   if (s->io.read) {
1702+      if (!(s->io.eof)(s->io_user_data)) return 0;
1703+      /*
1704+       * if feof() is true, check if buffer = end
1705+       * special case: we've only got the special 0 character at the end
1706+       */
1707+      if (s->read_from_callbacks == 0) return 1;
1708+   }
1709+
1710+   return s->img_buffer >= s->img_buffer_end;
1711+}
1712+#endif
1713+
1714+#if defined(STBI_NO_JPEG) && defined(STBI_NO_PNG) && defined(STBI_NO_BMP) && defined(STBI_NO_PSD) && defined(STBI_NO_TGA) && defined(STBI_NO_GIF) && defined(STBI_NO_PIC)
1715+/* nothing */
1716+#else
1717+static void stbi__skip(stbi__context *s, int n)
1718+{
1719+   if (n == 0) return; /* already there! */
1720+   if (n < 0) {
1721+      s->img_buffer = s->img_buffer_end;
1722+      return;
1723+   }
1724+   if (s->io.read) {
1725+      int blen = (int) (s->img_buffer_end - s->img_buffer);
1726+      if (blen < n) {
1727+         s->img_buffer = s->img_buffer_end;
1728+         (s->io.skip)(s->io_user_data, n - blen);
1729+         return;
1730+      }
1731+   }
1732+   s->img_buffer += n;
1733+}
1734+#endif
1735+
1736+#if defined(STBI_NO_PNG) && defined(STBI_NO_TGA) && defined(STBI_NO_HDR) && defined(STBI_NO_PNM)
1737+/* nothing */
1738+#else
1739+static int stbi__getn(stbi__context *s, stbi_uc *buffer, int n)
1740+{
1741+   if (s->io.read) {
1742+      int blen = (int) (s->img_buffer_end - s->img_buffer);
1743+      if (blen < n) {
1744+         int res, count;
1745+
1746+         memcpy(buffer, s->img_buffer, blen);
1747+
1748+         count = (s->io.read)(s->io_user_data, (char*) buffer + blen, n - blen);
1749+         res = (count == (n-blen));
1750+         s->img_buffer = s->img_buffer_end;
1751+         return res;
1752+      }
1753+   }
1754+
1755+   if (s->img_buffer+n <= s->img_buffer_end) {
1756+      memcpy(buffer, s->img_buffer, n);
1757+      s->img_buffer += n;
1758+      return 1;
1759+   } else
1760+      return 0;
1761+}
1762+#endif
1763+
1764+#if defined(STBI_NO_JPEG) && defined(STBI_NO_PNG) && defined(STBI_NO_PSD) && defined(STBI_NO_PIC)
1765+/* nothing */
1766+#else
1767+static int stbi__get16be(stbi__context *s)
1768+{
1769+   int z = stbi__get8(s);
1770+   return (z << 8) + stbi__get8(s);
1771+}
1772+#endif
1773+
1774+#if defined(STBI_NO_PNG) && defined(STBI_NO_PSD) && defined(STBI_NO_PIC)
1775+/* nothing */
1776+#else
1777+static stbi__uint32 stbi__get32be(stbi__context *s)
1778+{
1779+   stbi__uint32 z = stbi__get16be(s);
1780+   return (z << 16) + stbi__get16be(s);
1781+}
1782+#endif
1783+
1784+#if defined(STBI_NO_BMP) && defined(STBI_NO_TGA) && defined(STBI_NO_GIF)
1785+/* nothing */
1786+#else
1787+static int stbi__get16le(stbi__context *s)
1788+{
1789+   int z = stbi__get8(s);
1790+   return z + (stbi__get8(s) << 8);
1791+}
1792+#endif
1793+
1794+#ifndef STBI_NO_BMP
1795+static stbi__uint32 stbi__get32le(stbi__context *s)
1796+{
1797+   stbi__uint32 z = stbi__get16le(s);
1798+   z += (stbi__uint32)stbi__get16le(s) << 16;
1799+   return z;
1800+}
1801+#endif
1802+
1803+#define STBI__BYTECAST(x)  ((stbi_uc) ((x) & 255)) /* truncate int to byte without warnings */
1804+
1805+#if defined(STBI_NO_JPEG) && defined(STBI_NO_PNG) && defined(STBI_NO_BMP) && defined(STBI_NO_PSD) && defined(STBI_NO_TGA) && defined(STBI_NO_GIF) && defined(STBI_NO_PIC) && defined(STBI_NO_PNM)
1806+/* nothing */
1807+#else
1808+/*////////////////////////////////////////////////////////////////////////////*/
1809+/*
1810+ *
1811+ *  generic converter from built-in img_n to req_comp
1812+ *    individual types do this automatically as much as possible (e.g. jpeg
1813+ *    does all cases internally since it needs to colorspace convert anyway,
1814+ *    and it never has alpha, so very few cases ). png can automatically
1815+ *    interleave an alpha=255 channel, but falls back to this for other cases
1816+ *
1817+ *  assume data buffer is malloced, so malloc a new one and free that one
1818+ *  only failure mode is malloc failing
1819+ */
1820+
1821+static stbi_uc stbi__compute_y(int r, int g, int b)
1822+{
1823+   return (stbi_uc) (((r*77) + (g*150) +  (29*b)) >> 8);
1824+}
1825+#endif
1826+
1827+#if defined(STBI_NO_PNG) && defined(STBI_NO_BMP) && defined(STBI_NO_PSD) && defined(STBI_NO_TGA) && defined(STBI_NO_GIF) && defined(STBI_NO_PIC) && defined(STBI_NO_PNM)
1828+/* nothing */
1829+#else
1830+static unsigned char *stbi__convert_format(unsigned char *data, int img_n, int req_comp, unsigned int x, unsigned int y)
1831+{
1832+   int i,j;
1833+   unsigned char *good;
1834+
1835+   if (req_comp == img_n) return data;
1836+   STBI_ASSERT(req_comp >= 1 && req_comp <= 4);
1837+
1838+   good = (unsigned char *) stbi__malloc_mad3(req_comp, x, y, 0);
1839+   if (good == NULL) {
1840+      STBI_FREE(data);
1841+      return stbi__errpuc("outofmem", "Out of memory");
1842+   }
1843+
1844+   for (j=0; j < (int) y; ++j) {
1845+      unsigned char *src  = data + j * x * img_n   ;
1846+      unsigned char *dest = good + j * x * req_comp;
1847+
1848+      #define STBI__COMBO(a,b)  ((a)*8+(b))
1849+      #define STBI__CASE(a,b)   case STBI__COMBO(a,b): for(i=x-1; i >= 0; --i, src += a, dest += b)
1850+      /*
1851+       * convert source image with img_n components to one with req_comp components;
1852+       * avoid switch per pixel, so use switch per scanline and massive macros
1853+       */
1854+      switch (STBI__COMBO(img_n, req_comp)) {
1855+         STBI__CASE(1,2) { dest[0]=src[0]; dest[1]=255;                                     } break;
1856+         STBI__CASE(1,3) { dest[0]=dest[1]=dest[2]=src[0];                                  } break;
1857+         STBI__CASE(1,4) { dest[0]=dest[1]=dest[2]=src[0]; dest[3]=255;                     } break;
1858+         STBI__CASE(2,1) { dest[0]=src[0];                                                  } break;
1859+         STBI__CASE(2,3) { dest[0]=dest[1]=dest[2]=src[0];                                  } break;
1860+         STBI__CASE(2,4) { dest[0]=dest[1]=dest[2]=src[0]; dest[3]=src[1];                  } break;
1861+         STBI__CASE(3,4) { dest[0]=src[0];dest[1]=src[1];dest[2]=src[2];dest[3]=255;        } break;
1862+         STBI__CASE(3,1) { dest[0]=stbi__compute_y(src[0],src[1],src[2]);                   } break;
1863+         STBI__CASE(3,2) { dest[0]=stbi__compute_y(src[0],src[1],src[2]); dest[1] = 255;    } break;
1864+         STBI__CASE(4,1) { dest[0]=stbi__compute_y(src[0],src[1],src[2]);                   } break;
1865+         STBI__CASE(4,2) { dest[0]=stbi__compute_y(src[0],src[1],src[2]); dest[1] = src[3]; } break;
1866+         STBI__CASE(4,3) { dest[0]=src[0];dest[1]=src[1];dest[2]=src[2];                    } break;
1867+         default: STBI_ASSERT(0); STBI_FREE(data); STBI_FREE(good); return stbi__errpuc("unsupported", "Unsupported format conversion");
1868+      }
1869+      #undef STBI__CASE
1870+   }
1871+
1872+   STBI_FREE(data);
1873+   return good;
1874+}
1875+#endif
1876+
1877+#if defined(STBI_NO_PNG) && defined(STBI_NO_PSD)
1878+/* nothing */
1879+#else
1880+static stbi__uint16 stbi__compute_y_16(int r, int g, int b)
1881+{
1882+   return (stbi__uint16) (((r*77) + (g*150) +  (29*b)) >> 8);
1883+}
1884+#endif
1885+
1886+#if defined(STBI_NO_PNG) && defined(STBI_NO_PSD)
1887+/* nothing */
1888+#else
1889+static stbi__uint16 *stbi__convert_format16(stbi__uint16 *data, int img_n, int req_comp, unsigned int x, unsigned int y)
1890+{
1891+   int i,j;
1892+   stbi__uint16 *good;
1893+
1894+   if (req_comp == img_n) return data;
1895+   STBI_ASSERT(req_comp >= 1 && req_comp <= 4);
1896+
1897+   good = (stbi__uint16 *) stbi__malloc(req_comp * x * y * 2);
1898+   if (good == NULL) {
1899+      STBI_FREE(data);
1900+      return (stbi__uint16 *) stbi__errpuc("outofmem", "Out of memory");
1901+   }
1902+
1903+   for (j=0; j < (int) y; ++j) {
1904+      stbi__uint16 *src  = data + j * x * img_n   ;
1905+      stbi__uint16 *dest = good + j * x * req_comp;
1906+
1907+      #define STBI__COMBO(a,b)  ((a)*8+(b))
1908+      #define STBI__CASE(a,b)   case STBI__COMBO(a,b): for(i=x-1; i >= 0; --i, src += a, dest += b)
1909+      /*
1910+       * convert source image with img_n components to one with req_comp components;
1911+       * avoid switch per pixel, so use switch per scanline and massive macros
1912+       */
1913+      switch (STBI__COMBO(img_n, req_comp)) {
1914+         STBI__CASE(1,2) { dest[0]=src[0]; dest[1]=0xffff;                                     } break;
1915+         STBI__CASE(1,3) { dest[0]=dest[1]=dest[2]=src[0];                                     } break;
1916+         STBI__CASE(1,4) { dest[0]=dest[1]=dest[2]=src[0]; dest[3]=0xffff;                     } break;
1917+         STBI__CASE(2,1) { dest[0]=src[0];                                                     } break;
1918+         STBI__CASE(2,3) { dest[0]=dest[1]=dest[2]=src[0];                                     } break;
1919+         STBI__CASE(2,4) { dest[0]=dest[1]=dest[2]=src[0]; dest[3]=src[1];                     } break;
1920+         STBI__CASE(3,4) { dest[0]=src[0];dest[1]=src[1];dest[2]=src[2];dest[3]=0xffff;        } break;
1921+         STBI__CASE(3,1) { dest[0]=stbi__compute_y_16(src[0],src[1],src[2]);                   } break;
1922+         STBI__CASE(3,2) { dest[0]=stbi__compute_y_16(src[0],src[1],src[2]); dest[1] = 0xffff; } break;
1923+         STBI__CASE(4,1) { dest[0]=stbi__compute_y_16(src[0],src[1],src[2]);                   } break;
1924+         STBI__CASE(4,2) { dest[0]=stbi__compute_y_16(src[0],src[1],src[2]); dest[1] = src[3]; } break;
1925+         STBI__CASE(4,3) { dest[0]=src[0];dest[1]=src[1];dest[2]=src[2];                       } break;
1926+         default: STBI_ASSERT(0); STBI_FREE(data); STBI_FREE(good); return (stbi__uint16*) stbi__errpuc("unsupported", "Unsupported format conversion");
1927+      }
1928+      #undef STBI__CASE
1929+   }
1930+
1931+   STBI_FREE(data);
1932+   return good;
1933+}
1934+#endif
1935+
1936+#ifndef STBI_NO_LINEAR
1937+static float   *stbi__ldr_to_hdr(stbi_uc *data, int x, int y, int comp)
1938+{
1939+   int i,k,n;
1940+   float *output;
1941+   if (!data) return NULL;
1942+   output = (float *) stbi__malloc_mad4(x, y, comp, sizeof(float), 0);
1943+   if (output == NULL) { STBI_FREE(data); return stbi__errpf("outofmem", "Out of memory"); }
1944+   /* compute number of non-alpha components */
1945+   if (comp & 1) n = comp; else n = comp-1;
1946+   for (i=0; i < x*y; ++i) {
1947+      for (k=0; k < n; ++k) {
1948+         output[i*comp + k] = (float) (pow(data[i*comp+k]/255.0f, stbi__l2h_gamma) * stbi__l2h_scale);
1949+      }
1950+   }
1951+   if (n < comp) {
1952+      for (i=0; i < x*y; ++i) {
1953+         output[i*comp + n] = data[i*comp + n]/255.0f;
1954+      }
1955+   }
1956+   STBI_FREE(data);
1957+   return output;
1958+}
1959+#endif
1960+
1961+#ifndef STBI_NO_HDR
1962+#define stbi__float2int(x)   ((int) (x))
1963+static stbi_uc *stbi__hdr_to_ldr(float   *data, int x, int y, int comp)
1964+{
1965+   int i,k,n;
1966+   stbi_uc *output;
1967+   if (!data) return NULL;
1968+   output = (stbi_uc *) stbi__malloc_mad3(x, y, comp, 0);
1969+   if (output == NULL) { STBI_FREE(data); return stbi__errpuc("outofmem", "Out of memory"); }
1970+   /* compute number of non-alpha components */
1971+   if (comp & 1) n = comp; else n = comp-1;
1972+   for (i=0; i < x*y; ++i) {
1973+      for (k=0; k < n; ++k) {
1974+         float z = (float) pow(data[i*comp+k]*stbi__h2l_scale_i, stbi__h2l_gamma_i) * 255 + 0.5f;
1975+         if (z < 0) z = 0;
1976+         if (z > 255) z = 255;
1977+         output[i*comp + k] = (stbi_uc) stbi__float2int(z);
1978+      }
1979+      if (k < comp) {
1980+         float z = data[i*comp+k] * 255 + 0.5f;
1981+         if (z < 0) z = 0;
1982+         if (z > 255) z = 255;
1983+         output[i*comp + k] = (stbi_uc) stbi__float2int(z);
1984+      }
1985+   }
1986+   STBI_FREE(data);
1987+   return output;
1988+}
1989+#endif
1990+
1991+/*////////////////////////////////////////////////////////////////////////////*/
1992+/*
1993+ *
1994+ *  "baseline" JPEG/JFIF decoder
1995+ *
1996+ *    simple implementation
1997+ *      - doesn't support delayed output of y-dimension
1998+ *      - simple interface (only one output format: 8-bit interleaved RGB)
1999+ *      - doesn't try to recover corrupt jpegs
2000+ *      - doesn't allow partial loading, loading multiple at once
2001+ *      - still fast on x86 (copying globals into locals doesn't help x86)
2002+ *      - allocates lots of intermediate memory (full size of all components)
2003+ *        - non-interleaved case requires this anyway
2004+ *        - allows good upsampling (see next)
2005+ *    high-quality
2006+ *      - upsampled channels are bilinearly interpolated, even across blocks
2007+ *      - quality integer IDCT derived from IJG's 'slow'
2008+ *    performance
2009+ *      - fast huffman; reasonable integer IDCT
2010+ *      - some SIMD kernels for common paths on targets with SSE2/NEON
2011+ *      - uses a lot of intermediate memory, could cache poorly
2012+ */
2013+
2014+#ifndef STBI_NO_JPEG
2015+
2016+/* huffman decoding acceleration */
2017+#define FAST_BITS   9 /* larger handles more cases; smaller stomps less cache */
2018+
2019+typedef struct
2020+{
2021+   stbi_uc  fast[1 << FAST_BITS];
2022+   /* weirdly, repacking this into AoS is a 10% speed loss, instead of a win */
2023+   stbi__uint16 code[256];
2024+   stbi_uc  values[256];
2025+   stbi_uc  size[257];
2026+   unsigned int maxcode[18];
2027+   int    delta[17]; /* old 'firstsymbol' - old 'firstcode' */
2028+} stbi__huffman;
2029+
2030+typedef struct
2031+{
2032+   stbi__context *s;
2033+   stbi__huffman huff_dc[4];
2034+   stbi__huffman huff_ac[4];
2035+   stbi__uint16 dequant[4][64];
2036+   stbi__int16 fast_ac[4][1 << FAST_BITS];
2037+
2038+/* sizes for components, interleaved MCUs */
2039+   int img_h_max, img_v_max;
2040+   int img_mcu_x, img_mcu_y;
2041+   int img_mcu_w, img_mcu_h;
2042+
2043+/* definition of jpeg image component */
2044+   struct
2045+   {
2046+      int id;
2047+      int h,v;
2048+      int tq;
2049+      int hd,ha;
2050+      int dc_pred;
2051+
2052+      int x,y,w2,h2;
2053+      stbi_uc *data;
2054+      void *raw_data, *raw_coeff;
2055+      stbi_uc *linebuf;
2056+      short   *coeff; /* progressive only */
2057+      int      coeff_w, coeff_h; /* number of 8x8 coefficient blocks */
2058+   } img_comp[4];
2059+
2060+   stbi__uint32   code_buffer; /* jpeg entropy-coded buffer */
2061+   int            code_bits; /* number of valid bits */
2062+   unsigned char  marker; /* marker seen while filling entropy buffer */
2063+   int            nomore; /* flag if we saw a marker so must stop */
2064+
2065+   int            progressive;
2066+   int            spec_start;
2067+   int            spec_end;
2068+   int            succ_high;
2069+   int            succ_low;
2070+   int            eob_run;
2071+   int            jfif;
2072+   int            app14_color_transform; /* Adobe APP14 tag */
2073+   int            rgb;
2074+
2075+   int scan_n, order[4];
2076+   int restart_interval, todo;
2077+
2078+/* kernels */
2079+   void (*idct_block_kernel)(stbi_uc *out, int out_stride, short data[64]);
2080+   void (*YCbCr_to_RGB_kernel)(stbi_uc *out, const stbi_uc *y, const stbi_uc *pcb, const stbi_uc *pcr, int count, int step);
2081+   stbi_uc *(*resample_row_hv_2_kernel)(stbi_uc *out, stbi_uc *in_near, stbi_uc *in_far, int w, int hs);
2082+} stbi__jpeg;
2083+
2084+static int stbi__build_huffman(stbi__huffman *h, int *count)
2085+{
2086+   int i,j,k=0;
2087+   unsigned int code;
2088+   /* build size list for each symbol (from JPEG spec) */
2089+   for (i=0; i < 16; ++i) {
2090+      for (j=0; j < count[i]; ++j) {
2091+         h->size[k++] = (stbi_uc) (i+1);
2092+         if(k >= 257) return stbi__err("bad size list","Corrupt JPEG");
2093+      }
2094+   }
2095+   h->size[k] = 0;
2096+
2097+   /* compute actual symbols (from jpeg spec) */
2098+   code = 0;
2099+   k = 0;
2100+   for(j=1; j <= 16; ++j) {
2101+      /* compute delta to add to code to compute symbol id */
2102+      h->delta[j] = k - code;
2103+      if (h->size[k] == j) {
2104+         while (h->size[k] == j)
2105+            h->code[k++] = (stbi__uint16) (code++);
2106+         if (code-1 >= (1u << j)) return stbi__err("bad code lengths","Corrupt JPEG");
2107+      }
2108+      /* compute largest code + 1 for this size, preshifted as needed later */
2109+      h->maxcode[j] = code << (16-j);
2110+      code <<= 1;
2111+   }
2112+   h->maxcode[j] = 0xffffffff;
2113+
2114+   /* build non-spec acceleration table; 255 is flag for not-accelerated */
2115+   memset(h->fast, 255, 1 << FAST_BITS);
2116+   for (i=0; i < k; ++i) {
2117+      int s = h->size[i];
2118+      if (s <= FAST_BITS) {
2119+         int c = h->code[i] << (FAST_BITS-s);
2120+         int m = 1 << (FAST_BITS-s);
2121+         for (j=0; j < m; ++j) {
2122+            h->fast[c+j] = (stbi_uc) i;
2123+         }
2124+      }
2125+   }
2126+   return 1;
2127+}
2128+
2129+/*
2130+ * build a table that decodes both magnitude and value of small ACs in
2131+ * one go.
2132+ */
2133+static void stbi__build_fast_ac(stbi__int16 *fast_ac, stbi__huffman *h)
2134+{
2135+   int i;
2136+   for (i=0; i < (1 << FAST_BITS); ++i) {
2137+      stbi_uc fast = h->fast[i];
2138+      fast_ac[i] = 0;
2139+      if (fast < 255) {
2140+         int rs = h->values[fast];
2141+         int run = (rs >> 4) & 15;
2142+         int magbits = rs & 15;
2143+         int len = h->size[fast];
2144+
2145+         if (magbits && len + magbits <= FAST_BITS) {
2146+            /* magnitude code followed by receive_extend code */
2147+            int k = ((i << len) & ((1 << FAST_BITS) - 1)) >> (FAST_BITS - magbits);
2148+            int m = 1 << (magbits - 1);
2149+            if (k < m) k += (~0U << magbits) + 1;
2150+            /* if the result is small enough, we can fit it in fast_ac table */
2151+            if (k >= -128 && k <= 127)
2152+               fast_ac[i] = (stbi__int16) ((k * 256) + (run * 16) + (len + magbits));
2153+         }
2154+      }
2155+   }
2156+}
2157+
2158+static void stbi__grow_buffer_unsafe(stbi__jpeg *j)
2159+{
2160+   do {
2161+      unsigned int b = j->nomore ? 0 : stbi__get8(j->s);
2162+      if (b == 0xff) {
2163+         int c = stbi__get8(j->s);
2164+         while (c == 0xff) c = stbi__get8(j->s); /* consume fill bytes */
2165+         if (c != 0) {
2166+            j->marker = (unsigned char) c;
2167+            j->nomore = 1;
2168+            return;
2169+         }
2170+      }
2171+      j->code_buffer |= b << (24 - j->code_bits);
2172+      j->code_bits += 8;
2173+   } while (j->code_bits <= 24);
2174+}
2175+
2176+/* (1 << n) - 1 */
2177+static const stbi__uint32 stbi__bmask[17]={0,1,3,7,15,31,63,127,255,511,1023,2047,4095,8191,16383,32767,65535};
2178+
2179+/* decode a jpeg huffman value from the bitstream */
2180+stbi_inline static int stbi__jpeg_huff_decode(stbi__jpeg *j, stbi__huffman *h)
2181+{
2182+   unsigned int temp;
2183+   int c,k;
2184+
2185+   if (j->code_bits < 16) stbi__grow_buffer_unsafe(j);
2186+
2187+   /*
2188+    * look at the top FAST_BITS and determine what symbol ID it is,
2189+    * if the code is <= FAST_BITS
2190+    */
2191+   c = (j->code_buffer >> (32 - FAST_BITS)) & ((1 << FAST_BITS)-1);
2192+   k = h->fast[c];
2193+   if (k < 255) {
2194+      int s = h->size[k];
2195+      if (s > j->code_bits)
2196+         return -1;
2197+      j->code_buffer <<= s;
2198+      j->code_bits -= s;
2199+      return h->values[k];
2200+   }
2201+
2202+   /*
2203+    * naive test is to shift the code_buffer down so k bits are
2204+    * valid, then test against maxcode. To speed this up, we've
2205+    * preshifted maxcode left so that it has (16-k) 0s at the
2206+    * end; in other words, regardless of the number of bits, it
2207+    * wants to be compared against something shifted to have 16;
2208+    * that way we don't need to shift inside the loop.
2209+    */
2210+   temp = j->code_buffer >> 16;
2211+   for (k=FAST_BITS+1 ; ; ++k)
2212+      if (temp < h->maxcode[k])
2213+         break;
2214+   if (k == 17) {
2215+      /* error! code not found */
2216+      j->code_bits -= 16;
2217+      return -1;
2218+   }
2219+
2220+   if (k > j->code_bits)
2221+      return -1;
2222+
2223+   /* convert the huffman code to the symbol id */
2224+   c = ((j->code_buffer >> (32 - k)) & stbi__bmask[k]) + h->delta[k];
2225+   if(c < 0 || c >= 256) /* symbol id out of bounds! */
2226+       return -1;
2227+   STBI_ASSERT((((j->code_buffer) >> (32 - h->size[c])) & stbi__bmask[h->size[c]]) == h->code[c]);
2228+
2229+   /* convert the id to a symbol */
2230+   j->code_bits -= k;
2231+   j->code_buffer <<= k;
2232+   return h->values[c];
2233+}
2234+
2235+/* bias[n] = (-1<<n) + 1 */
2236+static const int stbi__jbias[16] = {0,-1,-3,-7,-15,-31,-63,-127,-255,-511,-1023,-2047,-4095,-8191,-16383,-32767};
2237+
2238+/*
2239+ * combined JPEG 'receive' and JPEG 'extend', since baseline
2240+ * always extends everything it receives.
2241+ */
2242+stbi_inline static int stbi__extend_receive(stbi__jpeg *j, int n)
2243+{
2244+   unsigned int k;
2245+   int sgn;
2246+   if (j->code_bits < n) stbi__grow_buffer_unsafe(j);
2247+   if (j->code_bits < n) return 0; /* ran out of bits from stream, return 0s intead of continuing */
2248+
2249+   sgn = j->code_buffer >> 31; /* sign bit always in MSB; 0 if MSB clear (positive), 1 if MSB set (negative) */
2250+   k = stbi_lrot(j->code_buffer, n);
2251+   j->code_buffer = k & ~stbi__bmask[n];
2252+   k &= stbi__bmask[n];
2253+   j->code_bits -= n;
2254+   return k + (stbi__jbias[n] & (sgn - 1));
2255+}
2256+
2257+/* get some unsigned bits */
2258+stbi_inline static int stbi__jpeg_get_bits(stbi__jpeg *j, int n)
2259+{
2260+   unsigned int k;
2261+   if (j->code_bits < n) stbi__grow_buffer_unsafe(j);
2262+   if (j->code_bits < n) return 0; /* ran out of bits from stream, return 0s intead of continuing */
2263+   k = stbi_lrot(j->code_buffer, n);
2264+   j->code_buffer = k & ~stbi__bmask[n];
2265+   k &= stbi__bmask[n];
2266+   j->code_bits -= n;
2267+   return k;
2268+}
2269+
2270+stbi_inline static int stbi__jpeg_get_bit(stbi__jpeg *j)
2271+{
2272+   unsigned int k;
2273+   if (j->code_bits < 1) stbi__grow_buffer_unsafe(j);
2274+   if (j->code_bits < 1) return 0; /* ran out of bits from stream, return 0s intead of continuing */
2275+   k = j->code_buffer;
2276+   j->code_buffer <<= 1;
2277+   --j->code_bits;
2278+   return k & 0x80000000;
2279+}
2280+
2281+/*
2282+ * given a value that's at position X in the zigzag stream,
2283+ * where does it appear in the 8x8 matrix coded as row-major?
2284+ */
2285+static const stbi_uc stbi__jpeg_dezigzag[64+15] =
2286+{
2287+    0,  1,  8, 16,  9,  2,  3, 10,
2288+   17, 24, 32, 25, 18, 11,  4,  5,
2289+   12, 19, 26, 33, 40, 48, 41, 34,
2290+   27, 20, 13,  6,  7, 14, 21, 28,
2291+   35, 42, 49, 56, 57, 50, 43, 36,
2292+   29, 22, 15, 23, 30, 37, 44, 51,
2293+   58, 59, 52, 45, 38, 31, 39, 46,
2294+   53, 60, 61, 54, 47, 55, 62, 63,
2295+   /* let corrupt input sample past end */
2296+   63, 63, 63, 63, 63, 63, 63, 63,
2297+   63, 63, 63, 63, 63, 63, 63
2298+};
2299+
2300+/* decode one 64-entry block-- */
2301+static int stbi__jpeg_decode_block(stbi__jpeg *j, short data[64], stbi__huffman *hdc, stbi__huffman *hac, stbi__int16 *fac, int b, stbi__uint16 *dequant)
2302+{
2303+   int diff,dc,k;
2304+   int t;
2305+
2306+   if (j->code_bits < 16) stbi__grow_buffer_unsafe(j);
2307+   t = stbi__jpeg_huff_decode(j, hdc);
2308+   if (t < 0 || t > 15) return stbi__err("bad huffman code","Corrupt JPEG");
2309+
2310+   /* 0 all the ac values now so we can do it 32-bits at a time */
2311+   memset(data,0,64*sizeof(data[0]));
2312+
2313+   diff = t ? stbi__extend_receive(j, t) : 0;
2314+   if (!stbi__addints_valid(j->img_comp[b].dc_pred, diff)) return stbi__err("bad delta","Corrupt JPEG");
2315+   dc = j->img_comp[b].dc_pred + diff;
2316+   j->img_comp[b].dc_pred = dc;
2317+   if (!stbi__mul2shorts_valid(dc, dequant[0])) return stbi__err("can't merge dc and ac", "Corrupt JPEG");
2318+   data[0] = (short) (dc * dequant[0]);
2319+
2320+   /* decode AC components, see JPEG spec */
2321+   k = 1;
2322+   do {
2323+      unsigned int zig;
2324+      int c,r,s;
2325+      if (j->code_bits < 16) stbi__grow_buffer_unsafe(j);
2326+      c = (j->code_buffer >> (32 - FAST_BITS)) & ((1 << FAST_BITS)-1);
2327+      r = fac[c];
2328+      if (r) { /* fast-AC path */
2329+         k += (r >> 4) & 15; /* run */
2330+         s = r & 15; /* combined length */
2331+         if (s > j->code_bits) return stbi__err("bad huffman code", "Combined length longer than code bits available");
2332+         j->code_buffer <<= s;
2333+         j->code_bits -= s;
2334+         /* decode into unzigzag'd location */
2335+         zig = stbi__jpeg_dezigzag[k++];
2336+         data[zig] = (short) ((r >> 8) * dequant[zig]);
2337+      } else {
2338+         int rs = stbi__jpeg_huff_decode(j, hac);
2339+         if (rs < 0) return stbi__err("bad huffman code","Corrupt JPEG");
2340+         s = rs & 15;
2341+         r = rs >> 4;
2342+         if (s == 0) {
2343+            if (rs != 0xf0) break; /* end block */
2344+            k += 16;
2345+         } else {
2346+            k += r;
2347+            /* decode into unzigzag'd location */
2348+            zig = stbi__jpeg_dezigzag[k++];
2349+            data[zig] = (short) (stbi__extend_receive(j,s) * dequant[zig]);
2350+         }
2351+      }
2352+   } while (k < 64);
2353+   return 1;
2354+}
2355+
2356+static int stbi__jpeg_decode_block_prog_dc(stbi__jpeg *j, short data[64], stbi__huffman *hdc, int b)
2357+{
2358+   int diff,dc;
2359+   int t;
2360+   if (j->spec_end != 0) return stbi__err("can't merge dc and ac", "Corrupt JPEG");
2361+
2362+   if (j->code_bits < 16) stbi__grow_buffer_unsafe(j);
2363+
2364+   if (j->succ_high == 0) {
2365+      /* first scan for DC coefficient, must be first */
2366+      memset(data,0,64*sizeof(data[0])); /* 0 all the ac values now */
2367+      t = stbi__jpeg_huff_decode(j, hdc);
2368+      if (t < 0 || t > 15) return stbi__err("can't merge dc and ac", "Corrupt JPEG");
2369+      diff = t ? stbi__extend_receive(j, t) : 0;
2370+
2371+      if (!stbi__addints_valid(j->img_comp[b].dc_pred, diff)) return stbi__err("bad delta", "Corrupt JPEG");
2372+      dc = j->img_comp[b].dc_pred + diff;
2373+      j->img_comp[b].dc_pred = dc;
2374+      if (!stbi__mul2shorts_valid(dc, 1 << j->succ_low)) return stbi__err("can't merge dc and ac", "Corrupt JPEG");
2375+      data[0] = (short) (dc * (1 << j->succ_low));
2376+   } else {
2377+      /* refinement scan for DC coefficient */
2378+      if (stbi__jpeg_get_bit(j))
2379+         data[0] += (short) (1 << j->succ_low);
2380+   }
2381+   return 1;
2382+}
2383+
2384+/*
2385+ * @OPTIMIZE: store non-zigzagged during the decode passes,
2386+ * and only de-zigzag when dequantizing
2387+ */
2388+static int stbi__jpeg_decode_block_prog_ac(stbi__jpeg *j, short data[64], stbi__huffman *hac, stbi__int16 *fac)
2389+{
2390+   int k;
2391+   if (j->spec_start == 0) return stbi__err("can't merge dc and ac", "Corrupt JPEG");
2392+
2393+   if (j->succ_high == 0) {
2394+      int shift = j->succ_low;
2395+
2396+      if (j->eob_run) {
2397+         --j->eob_run;
2398+         return 1;
2399+      }
2400+
2401+      k = j->spec_start;
2402+      do {
2403+         unsigned int zig;
2404+         int c,r,s;
2405+         if (j->code_bits < 16) stbi__grow_buffer_unsafe(j);
2406+         c = (j->code_buffer >> (32 - FAST_BITS)) & ((1 << FAST_BITS)-1);
2407+         r = fac[c];
2408+         if (r) { /* fast-AC path */
2409+            k += (r >> 4) & 15; /* run */
2410+            s = r & 15; /* combined length */
2411+            if (s > j->code_bits) return stbi__err("bad huffman code", "Combined length longer than code bits available");
2412+            j->code_buffer <<= s;
2413+            j->code_bits -= s;
2414+            zig = stbi__jpeg_dezigzag[k++];
2415+            data[zig] = (short) ((r >> 8) * (1 << shift));
2416+         } else {
2417+            int rs = stbi__jpeg_huff_decode(j, hac);
2418+            if (rs < 0) return stbi__err("bad huffman code","Corrupt JPEG");
2419+            s = rs & 15;
2420+            r = rs >> 4;
2421+            if (s == 0) {
2422+               if (r < 15) {
2423+                  j->eob_run = (1 << r);
2424+                  if (r)
2425+                     j->eob_run += stbi__jpeg_get_bits(j, r);
2426+                  --j->eob_run;
2427+                  break;
2428+               }
2429+               k += 16;
2430+            } else {
2431+               k += r;
2432+               zig = stbi__jpeg_dezigzag[k++];
2433+               data[zig] = (short) (stbi__extend_receive(j,s) * (1 << shift));
2434+            }
2435+         }
2436+      } while (k <= j->spec_end);
2437+   } else {
2438+      /* refinement scan for these AC coefficients */
2439+
2440+      short bit = (short) (1 << j->succ_low);
2441+
2442+      if (j->eob_run) {
2443+         --j->eob_run;
2444+         for (k = j->spec_start; k <= j->spec_end; ++k) {
2445+            short *p = &data[stbi__jpeg_dezigzag[k]];
2446+            if (*p != 0)
2447+               if (stbi__jpeg_get_bit(j))
2448+                  if ((*p & bit)==0) {
2449+                     if (*p > 0)
2450+                        *p += bit;
2451+                     else
2452+                        *p -= bit;
2453+                  }
2454+         }
2455+      } else {
2456+         k = j->spec_start;
2457+         do {
2458+            int r,s;
2459+            int rs = stbi__jpeg_huff_decode(j, hac); /* @OPTIMIZE see if we can use the fast path here, advance-by-r is so slow, eh */
2460+            if (rs < 0) return stbi__err("bad huffman code","Corrupt JPEG");
2461+            s = rs & 15;
2462+            r = rs >> 4;
2463+            if (s == 0) {
2464+               if (r < 15) {
2465+                  j->eob_run = (1 << r) - 1;
2466+                  if (r)
2467+                     j->eob_run += stbi__jpeg_get_bits(j, r);
2468+                  r = 64; /* force end of block */
2469+               } else {
2470+                  /*
2471+                   * r=15 s=0 should write 16 0s, so we just do
2472+                   * a run of 15 0s and then write s (which is 0),
2473+                   * so we don't have to do anything special here
2474+                   */
2475+               }
2476+            } else {
2477+               if (s != 1) return stbi__err("bad huffman code", "Corrupt JPEG");
2478+               /* sign bit */
2479+               if (stbi__jpeg_get_bit(j))
2480+                  s = bit;
2481+               else
2482+                  s = -bit;
2483+            }
2484+
2485+            /* advance by r */
2486+            while (k <= j->spec_end) {
2487+               short *p = &data[stbi__jpeg_dezigzag[k++]];
2488+               if (*p != 0) {
2489+                  if (stbi__jpeg_get_bit(j))
2490+                     if ((*p & bit)==0) {
2491+                        if (*p > 0)
2492+                           *p += bit;
2493+                        else
2494+                           *p -= bit;
2495+                     }
2496+               } else {
2497+                  if (r == 0) {
2498+                     *p = (short) s;
2499+                     break;
2500+                  }
2501+                  --r;
2502+               }
2503+            }
2504+         } while (k <= j->spec_end);
2505+      }
2506+   }
2507+   return 1;
2508+}
2509+
2510+/* take a -128..127 value and stbi__clamp it and convert to 0..255 */
2511+stbi_inline static stbi_uc stbi__clamp(int x)
2512+{
2513+   /* trick to use a single test to catch both cases */
2514+   if ((unsigned int) x > 255) {
2515+      if (x < 0) return 0;
2516+      if (x > 255) return 255;
2517+   }
2518+   return (stbi_uc) x;
2519+}
2520+
2521+#define stbi__f2f(x)  ((int) (((x) * 4096 + 0.5)))
2522+#define stbi__fsh(x)  ((x) * 4096)
2523+
2524+/* derived from jidctint -- DCT_ISLOW */
2525+#define STBI__IDCT_1D(s0,s1,s2,s3,s4,s5,s6,s7) \
2526+   int t0,t1,t2,t3,p1,p2,p3,p4,p5,x0,x1,x2,x3; \
2527+   p2 = s2;                                    \
2528+   p3 = s6;                                    \
2529+   p1 = (p2+p3) * stbi__f2f(0.5411961f);       \
2530+   t2 = p1 + p3*stbi__f2f(-1.847759065f);      \
2531+   t3 = p1 + p2*stbi__f2f( 0.765366865f);      \
2532+   p2 = s0;                                    \
2533+   p3 = s4;                                    \
2534+   t0 = stbi__fsh(p2+p3);                      \
2535+   t1 = stbi__fsh(p2-p3);                      \
2536+   x0 = t0+t3;                                 \
2537+   x3 = t0-t3;                                 \
2538+   x1 = t1+t2;                                 \
2539+   x2 = t1-t2;                                 \
2540+   t0 = s7;                                    \
2541+   t1 = s5;                                    \
2542+   t2 = s3;                                    \
2543+   t3 = s1;                                    \
2544+   p3 = t0+t2;                                 \
2545+   p4 = t1+t3;                                 \
2546+   p1 = t0+t3;                                 \
2547+   p2 = t1+t2;                                 \
2548+   p5 = (p3+p4)*stbi__f2f( 1.175875602f);      \
2549+   t0 = t0*stbi__f2f( 0.298631336f);           \
2550+   t1 = t1*stbi__f2f( 2.053119869f);           \
2551+   t2 = t2*stbi__f2f( 3.072711026f);           \
2552+   t3 = t3*stbi__f2f( 1.501321110f);           \
2553+   p1 = p5 + p1*stbi__f2f(-0.899976223f);      \
2554+   p2 = p5 + p2*stbi__f2f(-2.562915447f);      \
2555+   p3 = p3*stbi__f2f(-1.961570560f);           \
2556+   p4 = p4*stbi__f2f(-0.390180644f);           \
2557+   t3 += p1+p4;                                \
2558+   t2 += p2+p3;                                \
2559+   t1 += p2+p4;                                \
2560+   t0 += p1+p3;
2561+
2562+static void stbi__idct_block(stbi_uc *out, int out_stride, short data[64])
2563+{
2564+   int i,val[64],*v=val;
2565+   stbi_uc *o;
2566+   short *d = data;
2567+
2568+   /* columns */
2569+   for (i=0; i < 8; ++i,++d, ++v) {
2570+      /* if all zeroes, shortcut -- this avoids dequantizing 0s and IDCTing */
2571+      if (d[ 8]==0 && d[16]==0 && d[24]==0 && d[32]==0
2572+           && d[40]==0 && d[48]==0 && d[56]==0) {
2573+         /*
2574+          *    no shortcut                 0     seconds
2575+          *    (1|2|3|4|5|6|7)==0          0     seconds
2576+          *    all separate               -0.047 seconds
2577+          *    1 && 2|3 && 4|5 && 6|7:    -0.047 seconds
2578+          */
2579+         int dcterm = d[0]*4;
2580+         v[0] = v[8] = v[16] = v[24] = v[32] = v[40] = v[48] = v[56] = dcterm;
2581+      } else {
2582+         STBI__IDCT_1D(d[ 0],d[ 8],d[16],d[24],d[32],d[40],d[48],d[56])
2583+         /*
2584+          * constants scaled things up by 1<<12; let's bring them back
2585+          * down, but keep 2 extra bits of precision
2586+          */
2587+         x0 += 512; x1 += 512; x2 += 512; x3 += 512;
2588+         v[ 0] = (x0+t3) >> 10;
2589+         v[56] = (x0-t3) >> 10;
2590+         v[ 8] = (x1+t2) >> 10;
2591+         v[48] = (x1-t2) >> 10;
2592+         v[16] = (x2+t1) >> 10;
2593+         v[40] = (x2-t1) >> 10;
2594+         v[24] = (x3+t0) >> 10;
2595+         v[32] = (x3-t0) >> 10;
2596+      }
2597+   }
2598+
2599+   for (i=0, v=val, o=out; i < 8; ++i,v+=8,o+=out_stride) {
2600+      /* no fast case since the first 1D IDCT spread components out */
2601+      STBI__IDCT_1D(v[0],v[1],v[2],v[3],v[4],v[5],v[6],v[7])
2602+      /*
2603+       * constants scaled things up by 1<<12, plus we had 1<<2 from first
2604+       * loop, plus horizontal and vertical each scale by sqrt(8) so together
2605+       * we've got an extra 1<<3, so 1<<17 total we need to remove.
2606+       * so we want to round that, which means adding 0.5 * 1<<17,
2607+       * aka 65536. Also, we'll end up with -128 to 127 that we want
2608+       * to encode as 0..255 by adding 128, so we'll add that before the shift
2609+       */
2610+      x0 += 65536 + (128<<17);
2611+      x1 += 65536 + (128<<17);
2612+      x2 += 65536 + (128<<17);
2613+      x3 += 65536 + (128<<17);
2614+      /*
2615+       * tried computing the shifts into temps, or'ing the temps to see
2616+       * if any were out of range, but that was slower
2617+       */
2618+      o[0] = stbi__clamp((x0+t3) >> 17);
2619+      o[7] = stbi__clamp((x0-t3) >> 17);
2620+      o[1] = stbi__clamp((x1+t2) >> 17);
2621+      o[6] = stbi__clamp((x1-t2) >> 17);
2622+      o[2] = stbi__clamp((x2+t1) >> 17);
2623+      o[5] = stbi__clamp((x2-t1) >> 17);
2624+      o[3] = stbi__clamp((x3+t0) >> 17);
2625+      o[4] = stbi__clamp((x3-t0) >> 17);
2626+   }
2627+}
2628+
2629+#ifdef STBI_SSE2
2630+/*
2631+ * sse2 integer IDCT. not the fastest possible implementation but it
2632+ * produces bit-identical results to the generic C version so it's
2633+ * fully "transparent".
2634+ */
2635+static void stbi__idct_simd(stbi_uc *out, int out_stride, short data[64])
2636+{
2637+   /* This is constructed to match our regular (generic) integer IDCT exactly. */
2638+   __m128i row0, row1, row2, row3, row4, row5, row6, row7;
2639+   __m128i tmp;
2640+
2641+   /* dot product constant: even elems=x, odd elems=y */
2642+   #define dct_const(x,y)  _mm_setr_epi16((x),(y),(x),(y),(x),(y),(x),(y))
2643+
2644+   /*
2645+    * out(0) = c0[even]*x + c0[odd]*y   (c0, x, y 16-bit, out 32-bit)
2646+    * out(1) = c1[even]*x + c1[odd]*y
2647+    */
2648+   #define dct_rot(out0,out1, x,y,c0,c1) \
2649+      __m128i c0##lo = _mm_unpacklo_epi16((x),(y)); \
2650+      __m128i c0##hi = _mm_unpackhi_epi16((x),(y)); \
2651+      __m128i out0##_l = _mm_madd_epi16(c0##lo, c0); \
2652+      __m128i out0##_h = _mm_madd_epi16(c0##hi, c0); \
2653+      __m128i out1##_l = _mm_madd_epi16(c0##lo, c1); \
2654+      __m128i out1##_h = _mm_madd_epi16(c0##hi, c1)
2655+
2656+   /* out = in << 12  (in 16-bit, out 32-bit) */
2657+   #define dct_widen(out, in) \
2658+      __m128i out##_l = _mm_srai_epi32(_mm_unpacklo_epi16(_mm_setzero_si128(), (in)), 4); \
2659+      __m128i out##_h = _mm_srai_epi32(_mm_unpackhi_epi16(_mm_setzero_si128(), (in)), 4)
2660+
2661+   /* wide add */
2662+   #define dct_wadd(out, a, b) \
2663+      __m128i out##_l = _mm_add_epi32(a##_l, b##_l); \
2664+      __m128i out##_h = _mm_add_epi32(a##_h, b##_h)
2665+
2666+   /* wide sub */
2667+   #define dct_wsub(out, a, b) \
2668+      __m128i out##_l = _mm_sub_epi32(a##_l, b##_l); \
2669+      __m128i out##_h = _mm_sub_epi32(a##_h, b##_h)
2670+
2671+   /* butterfly a/b, add bias, then shift by "s" and pack */
2672+   #define dct_bfly32o(out0, out1, a,b,bias,s) \
2673+      { \
2674+         __m128i abiased_l = _mm_add_epi32(a##_l, bias); \
2675+         __m128i abiased_h = _mm_add_epi32(a##_h, bias); \
2676+         dct_wadd(sum, abiased, b); \
2677+         dct_wsub(dif, abiased, b); \
2678+         out0 = _mm_packs_epi32(_mm_srai_epi32(sum_l, s), _mm_srai_epi32(sum_h, s)); \
2679+         out1 = _mm_packs_epi32(_mm_srai_epi32(dif_l, s), _mm_srai_epi32(dif_h, s)); \
2680+      }
2681+
2682+   /* 8-bit interleave step (for transposes) */
2683+   #define dct_interleave8(a, b) \
2684+      tmp = a; \
2685+      a = _mm_unpacklo_epi8(a, b); \
2686+      b = _mm_unpackhi_epi8(tmp, b)
2687+
2688+   /* 16-bit interleave step (for transposes) */
2689+   #define dct_interleave16(a, b) \
2690+      tmp = a; \
2691+      a = _mm_unpacklo_epi16(a, b); \
2692+      b = _mm_unpackhi_epi16(tmp, b)
2693+
2694+   #define dct_pass(bias,shift) \
2695+      { \
2696+         /* even part */ \
2697+         dct_rot(t2e,t3e, row2,row6, rot0_0,rot0_1); \
2698+         __m128i sum04 = _mm_add_epi16(row0, row4); \
2699+         __m128i dif04 = _mm_sub_epi16(row0, row4); \
2700+         dct_widen(t0e, sum04); \
2701+         dct_widen(t1e, dif04); \
2702+         dct_wadd(x0, t0e, t3e); \
2703+         dct_wsub(x3, t0e, t3e); \
2704+         dct_wadd(x1, t1e, t2e); \
2705+         dct_wsub(x2, t1e, t2e); \
2706+         /* odd part */ \
2707+         dct_rot(y0o,y2o, row7,row3, rot2_0,rot2_1); \
2708+         dct_rot(y1o,y3o, row5,row1, rot3_0,rot3_1); \
2709+         __m128i sum17 = _mm_add_epi16(row1, row7); \
2710+         __m128i sum35 = _mm_add_epi16(row3, row5); \
2711+         dct_rot(y4o,y5o, sum17,sum35, rot1_0,rot1_1); \
2712+         dct_wadd(x4, y0o, y4o); \
2713+         dct_wadd(x5, y1o, y5o); \
2714+         dct_wadd(x6, y2o, y5o); \
2715+         dct_wadd(x7, y3o, y4o); \
2716+         dct_bfly32o(row0,row7, x0,x7,bias,shift); \
2717+         dct_bfly32o(row1,row6, x1,x6,bias,shift); \
2718+         dct_bfly32o(row2,row5, x2,x5,bias,shift); \
2719+         dct_bfly32o(row3,row4, x3,x4,bias,shift); \
2720+      }
2721+
2722+   __m128i rot0_0 = dct_const(stbi__f2f(0.5411961f), stbi__f2f(0.5411961f) + stbi__f2f(-1.847759065f));
2723+   __m128i rot0_1 = dct_const(stbi__f2f(0.5411961f) + stbi__f2f( 0.765366865f), stbi__f2f(0.5411961f));
2724+   __m128i rot1_0 = dct_const(stbi__f2f(1.175875602f) + stbi__f2f(-0.899976223f), stbi__f2f(1.175875602f));
2725+   __m128i rot1_1 = dct_const(stbi__f2f(1.175875602f), stbi__f2f(1.175875602f) + stbi__f2f(-2.562915447f));
2726+   __m128i rot2_0 = dct_const(stbi__f2f(-1.961570560f) + stbi__f2f( 0.298631336f), stbi__f2f(-1.961570560f));
2727+   __m128i rot2_1 = dct_const(stbi__f2f(-1.961570560f), stbi__f2f(-1.961570560f) + stbi__f2f( 3.072711026f));
2728+   __m128i rot3_0 = dct_const(stbi__f2f(-0.390180644f) + stbi__f2f( 2.053119869f), stbi__f2f(-0.390180644f));
2729+   __m128i rot3_1 = dct_const(stbi__f2f(-0.390180644f), stbi__f2f(-0.390180644f) + stbi__f2f( 1.501321110f));
2730+
2731+   /* rounding biases in column/row passes, see stbi__idct_block for explanation. */
2732+   __m128i bias_0 = _mm_set1_epi32(512);
2733+   __m128i bias_1 = _mm_set1_epi32(65536 + (128<<17));
2734+
2735+   /* load */
2736+   row0 = _mm_load_si128((const __m128i *) (data + 0*8));
2737+   row1 = _mm_load_si128((const __m128i *) (data + 1*8));
2738+   row2 = _mm_load_si128((const __m128i *) (data + 2*8));
2739+   row3 = _mm_load_si128((const __m128i *) (data + 3*8));
2740+   row4 = _mm_load_si128((const __m128i *) (data + 4*8));
2741+   row5 = _mm_load_si128((const __m128i *) (data + 5*8));
2742+   row6 = _mm_load_si128((const __m128i *) (data + 6*8));
2743+   row7 = _mm_load_si128((const __m128i *) (data + 7*8));
2744+
2745+   /* column pass */
2746+   dct_pass(bias_0, 10);
2747+
2748+   {
2749+      /* 16bit 8x8 transpose pass 1 */
2750+      dct_interleave16(row0, row4);
2751+      dct_interleave16(row1, row5);
2752+      dct_interleave16(row2, row6);
2753+      dct_interleave16(row3, row7);
2754+
2755+      /* transpose pass 2 */
2756+      dct_interleave16(row0, row2);
2757+      dct_interleave16(row1, row3);
2758+      dct_interleave16(row4, row6);
2759+      dct_interleave16(row5, row7);
2760+
2761+      /* transpose pass 3 */
2762+      dct_interleave16(row0, row1);
2763+      dct_interleave16(row2, row3);
2764+      dct_interleave16(row4, row5);
2765+      dct_interleave16(row6, row7);
2766+   }
2767+
2768+   /* row pass */
2769+   dct_pass(bias_1, 17);
2770+
2771+   {
2772+      /* pack */
2773+      __m128i p0 = _mm_packus_epi16(row0, row1); /* a0a1a2a3...a7b0b1b2b3...b7 */
2774+      __m128i p1 = _mm_packus_epi16(row2, row3);
2775+      __m128i p2 = _mm_packus_epi16(row4, row5);
2776+      __m128i p3 = _mm_packus_epi16(row6, row7);
2777+
2778+      /* 8bit 8x8 transpose pass 1 */
2779+      dct_interleave8(p0, p2); /* a0e0a1e1... */
2780+      dct_interleave8(p1, p3); /* c0g0c1g1... */
2781+
2782+      /* transpose pass 2 */
2783+      dct_interleave8(p0, p1); /* a0c0e0g0... */
2784+      dct_interleave8(p2, p3); /* b0d0f0h0... */
2785+
2786+      /* transpose pass 3 */
2787+      dct_interleave8(p0, p2); /* a0b0c0d0... */
2788+      dct_interleave8(p1, p3); /* a4b4c4d4... */
2789+
2790+      /* store */
2791+      _mm_storel_epi64((__m128i *) out, p0); out += out_stride;
2792+      _mm_storel_epi64((__m128i *) out, _mm_shuffle_epi32(p0, 0x4e)); out += out_stride;
2793+      _mm_storel_epi64((__m128i *) out, p2); out += out_stride;
2794+      _mm_storel_epi64((__m128i *) out, _mm_shuffle_epi32(p2, 0x4e)); out += out_stride;
2795+      _mm_storel_epi64((__m128i *) out, p1); out += out_stride;
2796+      _mm_storel_epi64((__m128i *) out, _mm_shuffle_epi32(p1, 0x4e)); out += out_stride;
2797+      _mm_storel_epi64((__m128i *) out, p3); out += out_stride;
2798+      _mm_storel_epi64((__m128i *) out, _mm_shuffle_epi32(p3, 0x4e));
2799+   }
2800+
2801+#undef dct_const
2802+#undef dct_rot
2803+#undef dct_widen
2804+#undef dct_wadd
2805+#undef dct_wsub
2806+#undef dct_bfly32o
2807+#undef dct_interleave8
2808+#undef dct_interleave16
2809+#undef dct_pass
2810+}
2811+
2812+#endif /* STBI_SSE2 */
2813+
2814+#ifdef STBI_NEON
2815+
2816+/*
2817+ * NEON integer IDCT. should produce bit-identical
2818+ * results to the generic C version.
2819+ */
2820+static void stbi__idct_simd(stbi_uc *out, int out_stride, short data[64])
2821+{
2822+   int16x8_t row0, row1, row2, row3, row4, row5, row6, row7;
2823+
2824+   int16x4_t rot0_0 = vdup_n_s16(stbi__f2f(0.5411961f));
2825+   int16x4_t rot0_1 = vdup_n_s16(stbi__f2f(-1.847759065f));
2826+   int16x4_t rot0_2 = vdup_n_s16(stbi__f2f( 0.765366865f));
2827+   int16x4_t rot1_0 = vdup_n_s16(stbi__f2f( 1.175875602f));
2828+   int16x4_t rot1_1 = vdup_n_s16(stbi__f2f(-0.899976223f));
2829+   int16x4_t rot1_2 = vdup_n_s16(stbi__f2f(-2.562915447f));
2830+   int16x4_t rot2_0 = vdup_n_s16(stbi__f2f(-1.961570560f));
2831+   int16x4_t rot2_1 = vdup_n_s16(stbi__f2f(-0.390180644f));
2832+   int16x4_t rot3_0 = vdup_n_s16(stbi__f2f( 0.298631336f));
2833+   int16x4_t rot3_1 = vdup_n_s16(stbi__f2f( 2.053119869f));
2834+   int16x4_t rot3_2 = vdup_n_s16(stbi__f2f( 3.072711026f));
2835+   int16x4_t rot3_3 = vdup_n_s16(stbi__f2f( 1.501321110f));
2836+
2837+#define dct_long_mul(out, inq, coeff) \
2838+   int32x4_t out##_l = vmull_s16(vget_low_s16(inq), coeff); \
2839+   int32x4_t out##_h = vmull_s16(vget_high_s16(inq), coeff)
2840+
2841+#define dct_long_mac(out, acc, inq, coeff) \
2842+   int32x4_t out##_l = vmlal_s16(acc##_l, vget_low_s16(inq), coeff); \
2843+   int32x4_t out##_h = vmlal_s16(acc##_h, vget_high_s16(inq), coeff)
2844+
2845+#define dct_widen(out, inq) \
2846+   int32x4_t out##_l = vshll_n_s16(vget_low_s16(inq), 12); \
2847+   int32x4_t out##_h = vshll_n_s16(vget_high_s16(inq), 12)
2848+
2849+/* wide add */
2850+#define dct_wadd(out, a, b) \
2851+   int32x4_t out##_l = vaddq_s32(a##_l, b##_l); \
2852+   int32x4_t out##_h = vaddq_s32(a##_h, b##_h)
2853+
2854+/* wide sub */
2855+#define dct_wsub(out, a, b) \
2856+   int32x4_t out##_l = vsubq_s32(a##_l, b##_l); \
2857+   int32x4_t out##_h = vsubq_s32(a##_h, b##_h)
2858+
2859+/* butterfly a/b, then shift using "shiftop" by "s" and pack */
2860+#define dct_bfly32o(out0,out1, a,b,shiftop,s) \
2861+   { \
2862+      dct_wadd(sum, a, b); \
2863+      dct_wsub(dif, a, b); \
2864+      out0 = vcombine_s16(shiftop(sum_l, s), shiftop(sum_h, s)); \
2865+      out1 = vcombine_s16(shiftop(dif_l, s), shiftop(dif_h, s)); \
2866+   }
2867+
2868+#define dct_pass(shiftop, shift) \
2869+   { \
2870+      /* even part */ \
2871+      int16x8_t sum26 = vaddq_s16(row2, row6); \
2872+      dct_long_mul(p1e, sum26, rot0_0); \
2873+      dct_long_mac(t2e, p1e, row6, rot0_1); \
2874+      dct_long_mac(t3e, p1e, row2, rot0_2); \
2875+      int16x8_t sum04 = vaddq_s16(row0, row4); \
2876+      int16x8_t dif04 = vsubq_s16(row0, row4); \
2877+      dct_widen(t0e, sum04); \
2878+      dct_widen(t1e, dif04); \
2879+      dct_wadd(x0, t0e, t3e); \
2880+      dct_wsub(x3, t0e, t3e); \
2881+      dct_wadd(x1, t1e, t2e); \
2882+      dct_wsub(x2, t1e, t2e); \
2883+      /* odd part */ \
2884+      int16x8_t sum15 = vaddq_s16(row1, row5); \
2885+      int16x8_t sum17 = vaddq_s16(row1, row7); \
2886+      int16x8_t sum35 = vaddq_s16(row3, row5); \
2887+      int16x8_t sum37 = vaddq_s16(row3, row7); \
2888+      int16x8_t sumodd = vaddq_s16(sum17, sum35); \
2889+      dct_long_mul(p5o, sumodd, rot1_0); \
2890+      dct_long_mac(p1o, p5o, sum17, rot1_1); \
2891+      dct_long_mac(p2o, p5o, sum35, rot1_2); \
2892+      dct_long_mul(p3o, sum37, rot2_0); \
2893+      dct_long_mul(p4o, sum15, rot2_1); \
2894+      dct_wadd(sump13o, p1o, p3o); \
2895+      dct_wadd(sump24o, p2o, p4o); \
2896+      dct_wadd(sump23o, p2o, p3o); \
2897+      dct_wadd(sump14o, p1o, p4o); \
2898+      dct_long_mac(x4, sump13o, row7, rot3_0); \
2899+      dct_long_mac(x5, sump24o, row5, rot3_1); \
2900+      dct_long_mac(x6, sump23o, row3, rot3_2); \
2901+      dct_long_mac(x7, sump14o, row1, rot3_3); \
2902+      dct_bfly32o(row0,row7, x0,x7,shiftop,shift); \
2903+      dct_bfly32o(row1,row6, x1,x6,shiftop,shift); \
2904+      dct_bfly32o(row2,row5, x2,x5,shiftop,shift); \
2905+      dct_bfly32o(row3,row4, x3,x4,shiftop,shift); \
2906+   }
2907+
2908+   /* load */
2909+   row0 = vld1q_s16(data + 0*8);
2910+   row1 = vld1q_s16(data + 1*8);
2911+   row2 = vld1q_s16(data + 2*8);
2912+   row3 = vld1q_s16(data + 3*8);
2913+   row4 = vld1q_s16(data + 4*8);
2914+   row5 = vld1q_s16(data + 5*8);
2915+   row6 = vld1q_s16(data + 6*8);
2916+   row7 = vld1q_s16(data + 7*8);
2917+
2918+   /* add DC bias */
2919+   row0 = vaddq_s16(row0, vsetq_lane_s16(1024, vdupq_n_s16(0), 0));
2920+
2921+   /* column pass */
2922+   dct_pass(vrshrn_n_s32, 10);
2923+
2924+   /* 16bit 8x8 transpose */
2925+   {
2926+/*
2927+ * these three map to a single VTRN.16, VTRN.32, and VSWP, respectively.
2928+ * whether compilers actually get this is another story, sadly.
2929+ */
2930+#define dct_trn16(x, y) { int16x8x2_t t = vtrnq_s16(x, y); x = t.val[0]; y = t.val[1]; }
2931+#define dct_trn32(x, y) { int32x4x2_t t = vtrnq_s32(vreinterpretq_s32_s16(x), vreinterpretq_s32_s16(y)); x = vreinterpretq_s16_s32(t.val[0]); y = vreinterpretq_s16_s32(t.val[1]); }
2932+#define dct_trn64(x, y) { int16x8_t x0 = x; int16x8_t y0 = y; x = vcombine_s16(vget_low_s16(x0), vget_low_s16(y0)); y = vcombine_s16(vget_high_s16(x0), vget_high_s16(y0)); }
2933+
2934+      /* pass 1 */
2935+      dct_trn16(row0, row1); /* a0b0a2b2a4b4a6b6 */
2936+      dct_trn16(row2, row3);
2937+      dct_trn16(row4, row5);
2938+      dct_trn16(row6, row7);
2939+
2940+      /* pass 2 */
2941+      dct_trn32(row0, row2); /* a0b0c0d0a4b4c4d4 */
2942+      dct_trn32(row1, row3);
2943+      dct_trn32(row4, row6);
2944+      dct_trn32(row5, row7);
2945+
2946+      /* pass 3 */
2947+      dct_trn64(row0, row4); /* a0b0c0d0e0f0g0h0 */
2948+      dct_trn64(row1, row5);
2949+      dct_trn64(row2, row6);
2950+      dct_trn64(row3, row7);
2951+
2952+#undef dct_trn16
2953+#undef dct_trn32
2954+#undef dct_trn64
2955+   }
2956+
2957+   /*
2958+    * row pass
2959+    * vrshrn_n_s32 only supports shifts up to 16, we need
2960+    * 17. so do a non-rounding shift of 16 first then follow
2961+    * up with a rounding shift by 1.
2962+    */
2963+   dct_pass(vshrn_n_s32, 16);
2964+
2965+   {
2966+      /* pack and round */
2967+      uint8x8_t p0 = vqrshrun_n_s16(row0, 1);
2968+      uint8x8_t p1 = vqrshrun_n_s16(row1, 1);
2969+      uint8x8_t p2 = vqrshrun_n_s16(row2, 1);
2970+      uint8x8_t p3 = vqrshrun_n_s16(row3, 1);
2971+      uint8x8_t p4 = vqrshrun_n_s16(row4, 1);
2972+      uint8x8_t p5 = vqrshrun_n_s16(row5, 1);
2973+      uint8x8_t p6 = vqrshrun_n_s16(row6, 1);
2974+      uint8x8_t p7 = vqrshrun_n_s16(row7, 1);
2975+
2976+      /* again, these can translate into one instruction, but often don't. */
2977+#define dct_trn8_8(x, y) { uint8x8x2_t t = vtrn_u8(x, y); x = t.val[0]; y = t.val[1]; }
2978+#define dct_trn8_16(x, y) { uint16x4x2_t t = vtrn_u16(vreinterpret_u16_u8(x), vreinterpret_u16_u8(y)); x = vreinterpret_u8_u16(t.val[0]); y = vreinterpret_u8_u16(t.val[1]); }
2979+#define dct_trn8_32(x, y) { uint32x2x2_t t = vtrn_u32(vreinterpret_u32_u8(x), vreinterpret_u32_u8(y)); x = vreinterpret_u8_u32(t.val[0]); y = vreinterpret_u8_u32(t.val[1]); }
2980+
2981+      /*
2982+       * sadly can't use interleaved stores here since we only write
2983+       * 8 bytes to each scan line!
2984+       */
2985+
2986+      /* 8x8 8-bit transpose pass 1 */
2987+      dct_trn8_8(p0, p1);
2988+      dct_trn8_8(p2, p3);
2989+      dct_trn8_8(p4, p5);
2990+      dct_trn8_8(p6, p7);
2991+
2992+      /* pass 2 */
2993+      dct_trn8_16(p0, p2);
2994+      dct_trn8_16(p1, p3);
2995+      dct_trn8_16(p4, p6);
2996+      dct_trn8_16(p5, p7);
2997+
2998+      /* pass 3 */
2999+      dct_trn8_32(p0, p4);
3000+      dct_trn8_32(p1, p5);
3001+      dct_trn8_32(p2, p6);
3002+      dct_trn8_32(p3, p7);
3003+
3004+      /* store */
3005+      vst1_u8(out, p0); out += out_stride;
3006+      vst1_u8(out, p1); out += out_stride;
3007+      vst1_u8(out, p2); out += out_stride;
3008+      vst1_u8(out, p3); out += out_stride;
3009+      vst1_u8(out, p4); out += out_stride;
3010+      vst1_u8(out, p5); out += out_stride;
3011+      vst1_u8(out, p6); out += out_stride;
3012+      vst1_u8(out, p7);
3013+
3014+#undef dct_trn8_8
3015+#undef dct_trn8_16
3016+#undef dct_trn8_32
3017+   }
3018+
3019+#undef dct_long_mul
3020+#undef dct_long_mac
3021+#undef dct_widen
3022+#undef dct_wadd
3023+#undef dct_wsub
3024+#undef dct_bfly32o
3025+#undef dct_pass
3026+}
3027+
3028+#endif /* STBI_NEON */
3029+
3030+#define STBI__MARKER_none  0xff
3031+/*
3032+ * if there's a pending marker from the entropy stream, return that
3033+ * otherwise, fetch from the stream and get a marker. if there's no
3034+ * marker, return 0xff, which is never a valid marker value
3035+ */
3036+static stbi_uc stbi__get_marker(stbi__jpeg *j)
3037+{
3038+   stbi_uc x;
3039+   if (j->marker != STBI__MARKER_none) { x = j->marker; j->marker = STBI__MARKER_none; return x; }
3040+   x = stbi__get8(j->s);
3041+   if (x != 0xff) return STBI__MARKER_none;
3042+   while (x == 0xff)
3043+      x = stbi__get8(j->s); /* consume repeated 0xff fill bytes */
3044+   return x;
3045+}
3046+
3047+/*
3048+ * in each scan, we'll have scan_n components, and the order
3049+ * of the components is specified by order[]
3050+ */
3051+#define STBI__RESTART(x)     ((x) >= 0xd0 && (x) <= 0xd7)
3052+
3053+/*
3054+ * after a restart interval, stbi__jpeg_reset the entropy decoder and
3055+ * the dc prediction
3056+ */
3057+static void stbi__jpeg_reset(stbi__jpeg *j)
3058+{
3059+   j->code_bits = 0;
3060+   j->code_buffer = 0;
3061+   j->nomore = 0;
3062+   j->img_comp[0].dc_pred = j->img_comp[1].dc_pred = j->img_comp[2].dc_pred = j->img_comp[3].dc_pred = 0;
3063+   j->marker = STBI__MARKER_none;
3064+   j->todo = j->restart_interval ? j->restart_interval : 0x7fffffff;
3065+   j->eob_run = 0;
3066+   /*
3067+    * no more than 1<<31 MCUs if no restart_interal? that's plenty safe,
3068+    * since we don't even allow 1<<30 pixels
3069+    */
3070+}
3071+
3072+static int stbi__parse_entropy_coded_data(stbi__jpeg *z)
3073+{
3074+   stbi__jpeg_reset(z);
3075+   if (!z->progressive) {
3076+      if (z->scan_n == 1) {
3077+         int i,j;
3078+         STBI_SIMD_ALIGN(short, data[64]);
3079+         int n = z->order[0];
3080+         /*
3081+          * non-interleaved data, we just need to process one block at a time,
3082+          * in trivial scanline order
3083+          * number of blocks to do just depends on how many actual "pixels" this
3084+          * component has, independent of interleaved MCU blocking and such
3085+          */
3086+         int w = (z->img_comp[n].x+7) >> 3;
3087+         int h = (z->img_comp[n].y+7) >> 3;
3088+         for (j=0; j < h; ++j) {
3089+            for (i=0; i < w; ++i) {
3090+               int ha = z->img_comp[n].ha;
3091+               if (!stbi__jpeg_decode_block(z, data, z->huff_dc+z->img_comp[n].hd, z->huff_ac+ha, z->fast_ac[ha], n, z->dequant[z->img_comp[n].tq])) return 0;
3092+               z->idct_block_kernel(z->img_comp[n].data+z->img_comp[n].w2*j*8+i*8, z->img_comp[n].w2, data);
3093+               /* every data block is an MCU, so countdown the restart interval */
3094+               if (--z->todo <= 0) {
3095+                  if (z->code_bits < 24) stbi__grow_buffer_unsafe(z);
3096+                  /*
3097+                   * if it's NOT a restart, then just bail, so we get corrupt data
3098+                   * rather than no data
3099+                   */
3100+                  if (!STBI__RESTART(z->marker)) return 1;
3101+                  stbi__jpeg_reset(z);
3102+               }
3103+            }
3104+         }
3105+         return 1;
3106+      } else { /* interleaved */
3107+         int i,j,k,x,y;
3108+         STBI_SIMD_ALIGN(short, data[64]);
3109+         for (j=0; j < z->img_mcu_y; ++j) {
3110+            for (i=0; i < z->img_mcu_x; ++i) {
3111+               /* scan an interleaved mcu... process scan_n components in order */
3112+               for (k=0; k < z->scan_n; ++k) {
3113+                  int n = z->order[k];
3114+                  /*
3115+                   * scan out an mcu's worth of this component; that's just determined
3116+                   * by the basic H and V specified for the component
3117+                   */
3118+                  for (y=0; y < z->img_comp[n].v; ++y) {
3119+                     for (x=0; x < z->img_comp[n].h; ++x) {
3120+                        int x2 = (i*z->img_comp[n].h + x)*8;
3121+                        int y2 = (j*z->img_comp[n].v + y)*8;
3122+                        int ha = z->img_comp[n].ha;
3123+                        if (!stbi__jpeg_decode_block(z, data, z->huff_dc+z->img_comp[n].hd, z->huff_ac+ha, z->fast_ac[ha], n, z->dequant[z->img_comp[n].tq])) return 0;
3124+                        z->idct_block_kernel(z->img_comp[n].data+z->img_comp[n].w2*y2+x2, z->img_comp[n].w2, data);
3125+                     }
3126+                  }
3127+               }
3128+               /*
3129+                * after all interleaved components, that's an interleaved MCU,
3130+                * so now count down the restart interval
3131+                */
3132+               if (--z->todo <= 0) {
3133+                  if (z->code_bits < 24) stbi__grow_buffer_unsafe(z);
3134+                  if (!STBI__RESTART(z->marker)) return 1;
3135+                  stbi__jpeg_reset(z);
3136+               }
3137+            }
3138+         }
3139+         return 1;
3140+      }
3141+   } else {
3142+      if (z->scan_n == 1) {
3143+         int i,j;
3144+         int n = z->order[0];
3145+         /*
3146+          * non-interleaved data, we just need to process one block at a time,
3147+          * in trivial scanline order
3148+          * number of blocks to do just depends on how many actual "pixels" this
3149+          * component has, independent of interleaved MCU blocking and such
3150+          */
3151+         int w = (z->img_comp[n].x+7) >> 3;
3152+         int h = (z->img_comp[n].y+7) >> 3;
3153+         for (j=0; j < h; ++j) {
3154+            for (i=0; i < w; ++i) {
3155+               short *data = z->img_comp[n].coeff + 64 * (i + j * z->img_comp[n].coeff_w);
3156+               if (z->spec_start == 0) {
3157+                  if (!stbi__jpeg_decode_block_prog_dc(z, data, &z->huff_dc[z->img_comp[n].hd], n))
3158+                     return 0;
3159+               } else {
3160+                  int ha = z->img_comp[n].ha;
3161+                  if (!stbi__jpeg_decode_block_prog_ac(z, data, &z->huff_ac[ha], z->fast_ac[ha]))
3162+                     return 0;
3163+               }
3164+               /* every data block is an MCU, so countdown the restart interval */
3165+               if (--z->todo <= 0) {
3166+                  if (z->code_bits < 24) stbi__grow_buffer_unsafe(z);
3167+                  if (!STBI__RESTART(z->marker)) return 1;
3168+                  stbi__jpeg_reset(z);
3169+               }
3170+            }
3171+         }
3172+         return 1;
3173+      } else { /* interleaved */
3174+         int i,j,k,x,y;
3175+         for (j=0; j < z->img_mcu_y; ++j) {
3176+            for (i=0; i < z->img_mcu_x; ++i) {
3177+               /* scan an interleaved mcu... process scan_n components in order */
3178+               for (k=0; k < z->scan_n; ++k) {
3179+                  int n = z->order[k];
3180+                  /*
3181+                   * scan out an mcu's worth of this component; that's just determined
3182+                   * by the basic H and V specified for the component
3183+                   */
3184+                  for (y=0; y < z->img_comp[n].v; ++y) {
3185+                     for (x=0; x < z->img_comp[n].h; ++x) {
3186+                        int x2 = (i*z->img_comp[n].h + x);
3187+                        int y2 = (j*z->img_comp[n].v + y);
3188+                        short *data = z->img_comp[n].coeff + 64 * (x2 + y2 * z->img_comp[n].coeff_w);
3189+                        if (!stbi__jpeg_decode_block_prog_dc(z, data, &z->huff_dc[z->img_comp[n].hd], n))
3190+                           return 0;
3191+                     }
3192+                  }
3193+               }
3194+               /*
3195+                * after all interleaved components, that's an interleaved MCU,
3196+                * so now count down the restart interval
3197+                */
3198+               if (--z->todo <= 0) {
3199+                  if (z->code_bits < 24) stbi__grow_buffer_unsafe(z);
3200+                  if (!STBI__RESTART(z->marker)) return 1;
3201+                  stbi__jpeg_reset(z);
3202+               }
3203+            }
3204+         }
3205+         return 1;
3206+      }
3207+   }
3208+}
3209+
3210+static void stbi__jpeg_dequantize(short *data, stbi__uint16 *dequant)
3211+{
3212+   int i;
3213+   for (i=0; i < 64; ++i)
3214+      data[i] *= dequant[i];
3215+}
3216+
3217+static void stbi__jpeg_finish(stbi__jpeg *z)
3218+{
3219+   if (z->progressive) {
3220+      /* dequantize and idct the data */
3221+      int i,j,n;
3222+      for (n=0; n < z->s->img_n; ++n) {
3223+         int w = (z->img_comp[n].x+7) >> 3;
3224+         int h = (z->img_comp[n].y+7) >> 3;
3225+         for (j=0; j < h; ++j) {
3226+            for (i=0; i < w; ++i) {
3227+               short *data = z->img_comp[n].coeff + 64 * (i + j * z->img_comp[n].coeff_w);
3228+               stbi__jpeg_dequantize(data, z->dequant[z->img_comp[n].tq]);
3229+               z->idct_block_kernel(z->img_comp[n].data+z->img_comp[n].w2*j*8+i*8, z->img_comp[n].w2, data);
3230+            }
3231+         }
3232+      }
3233+   }
3234+}
3235+
3236+static int stbi__process_marker(stbi__jpeg *z, int m)
3237+{
3238+   int L;
3239+   switch (m) {
3240+      case STBI__MARKER_none: /* no marker found */
3241+         return stbi__err("expected marker","Corrupt JPEG");
3242+
3243+      case 0xDD: /* DRI - specify restart interval */
3244+         if (stbi__get16be(z->s) != 4) return stbi__err("bad DRI len","Corrupt JPEG");
3245+         z->restart_interval = stbi__get16be(z->s);
3246+         return 1;
3247+
3248+      case 0xDB: /* DQT - define quantization table */
3249+         L = stbi__get16be(z->s)-2;
3250+         while (L > 0) {
3251+            int q = stbi__get8(z->s);
3252+            int p = q >> 4, sixteen = (p != 0);
3253+            int t = q & 15,i;
3254+            if (p != 0 && p != 1) return stbi__err("bad DQT type","Corrupt JPEG");
3255+            if (t > 3) return stbi__err("bad DQT table","Corrupt JPEG");
3256+
3257+            for (i=0; i < 64; ++i)
3258+               z->dequant[t][stbi__jpeg_dezigzag[i]] = (stbi__uint16)(sixteen ? stbi__get16be(z->s) : stbi__get8(z->s));
3259+            L -= (sixteen ? 129 : 65);
3260+         }
3261+         return L==0;
3262+
3263+      case 0xC4: /* DHT - define huffman table */
3264+         L = stbi__get16be(z->s)-2;
3265+         while (L > 0) {
3266+            stbi_uc *v;
3267+            int sizes[16],i,n=0;
3268+            int q = stbi__get8(z->s);
3269+            int tc = q >> 4;
3270+            int th = q & 15;
3271+            if (tc > 1 || th > 3) return stbi__err("bad DHT header","Corrupt JPEG");
3272+            for (i=0; i < 16; ++i) {
3273+               sizes[i] = stbi__get8(z->s);
3274+               n += sizes[i];
3275+            }
3276+            if(n > 256) return stbi__err("bad DHT header","Corrupt JPEG"); /* Loop over i < n would write past end of values! */
3277+            L -= 17;
3278+            if (tc == 0) {
3279+               if (!stbi__build_huffman(z->huff_dc+th, sizes)) return 0;
3280+               v = z->huff_dc[th].values;
3281+            } else {
3282+               if (!stbi__build_huffman(z->huff_ac+th, sizes)) return 0;
3283+               v = z->huff_ac[th].values;
3284+            }
3285+            for (i=0; i < n; ++i)
3286+               v[i] = stbi__get8(z->s);
3287+            if (tc != 0)
3288+               stbi__build_fast_ac(z->fast_ac[th], z->huff_ac + th);
3289+            L -= n;
3290+         }
3291+         return L==0;
3292+   }
3293+
3294+   /* check for comment block or APP blocks */
3295+   if ((m >= 0xE0 && m <= 0xEF) || m == 0xFE) {
3296+      L = stbi__get16be(z->s);
3297+      if (L < 2) {
3298+         if (m == 0xFE)
3299+            return stbi__err("bad COM len","Corrupt JPEG");
3300+         else
3301+            return stbi__err("bad APP len","Corrupt JPEG");
3302+      }
3303+      L -= 2;
3304+
3305+      if (m == 0xE0 && L >= 5) { /* JFIF APP0 segment */
3306+         static const unsigned char tag[5] = {'J','F','I','F','\0'};
3307+         int ok = 1;
3308+         int i;
3309+         for (i=0; i < 5; ++i)
3310+            if (stbi__get8(z->s) != tag[i])
3311+               ok = 0;
3312+         L -= 5;
3313+         if (ok)
3314+            z->jfif = 1;
3315+      } else if (m == 0xEE && L >= 12) { /* Adobe APP14 segment */
3316+         static const unsigned char tag[6] = {'A','d','o','b','e','\0'};
3317+         int ok = 1;
3318+         int i;
3319+         for (i=0; i < 6; ++i)
3320+            if (stbi__get8(z->s) != tag[i])
3321+               ok = 0;
3322+         L -= 6;
3323+         if (ok) {
3324+            stbi__get8(z->s); /* version */
3325+            stbi__get16be(z->s); /* flags0 */
3326+            stbi__get16be(z->s); /* flags1 */
3327+            z->app14_color_transform = stbi__get8(z->s); /* color transform */
3328+            L -= 6;
3329+         }
3330+      }
3331+
3332+      stbi__skip(z->s, L);
3333+      return 1;
3334+   }
3335+
3336+   return stbi__err("unknown marker","Corrupt JPEG");
3337+}
3338+
3339+/* after we see SOS */
3340+static int stbi__process_scan_header(stbi__jpeg *z)
3341+{
3342+   int i;
3343+   int Ls = stbi__get16be(z->s);
3344+   z->scan_n = stbi__get8(z->s);
3345+   if (z->scan_n < 1 || z->scan_n > 4 || z->scan_n > (int) z->s->img_n) return stbi__err("bad SOS component count","Corrupt JPEG");
3346+   if (Ls != 6+2*z->scan_n) return stbi__err("bad SOS len","Corrupt JPEG");
3347+   for (i=0; i < z->scan_n; ++i) {
3348+      int id = stbi__get8(z->s), which;
3349+      int q = stbi__get8(z->s);
3350+      for (which = 0; which < z->s->img_n; ++which)
3351+         if (z->img_comp[which].id == id)
3352+            break;
3353+      if (which == z->s->img_n) return 0; /* no match */
3354+      z->img_comp[which].hd = q >> 4;   if (z->img_comp[which].hd > 3) return stbi__err("bad DC huff","Corrupt JPEG");
3355+      z->img_comp[which].ha = q & 15;   if (z->img_comp[which].ha > 3) return stbi__err("bad AC huff","Corrupt JPEG");
3356+      z->order[i] = which;
3357+   }
3358+
3359+   {
3360+      int aa;
3361+      z->spec_start = stbi__get8(z->s);
3362+      z->spec_end   = stbi__get8(z->s); /* should be 63, but might be 0 */
3363+      aa = stbi__get8(z->s);
3364+      z->succ_high = (aa >> 4);
3365+      z->succ_low  = (aa & 15);
3366+      if (z->progressive) {
3367+         if (z->spec_start > 63 || z->spec_end > 63  || z->spec_start > z->spec_end || z->succ_high > 13 || z->succ_low > 13)
3368+            return stbi__err("bad SOS", "Corrupt JPEG");
3369+      } else {
3370+         if (z->spec_start != 0) return stbi__err("bad SOS","Corrupt JPEG");
3371+         if (z->succ_high != 0 || z->succ_low != 0) return stbi__err("bad SOS","Corrupt JPEG");
3372+         z->spec_end = 63;
3373+      }
3374+   }
3375+
3376+   return 1;
3377+}
3378+
3379+static int stbi__free_jpeg_components(stbi__jpeg *z, int ncomp, int why)
3380+{
3381+   int i;
3382+   for (i=0; i < ncomp; ++i) {
3383+      if (z->img_comp[i].raw_data) {
3384+         STBI_FREE(z->img_comp[i].raw_data);
3385+         z->img_comp[i].raw_data = NULL;
3386+         z->img_comp[i].data = NULL;
3387+      }
3388+      if (z->img_comp[i].raw_coeff) {
3389+         STBI_FREE(z->img_comp[i].raw_coeff);
3390+         z->img_comp[i].raw_coeff = 0;
3391+         z->img_comp[i].coeff = 0;
3392+      }
3393+      if (z->img_comp[i].linebuf) {
3394+         STBI_FREE(z->img_comp[i].linebuf);
3395+         z->img_comp[i].linebuf = NULL;
3396+      }
3397+   }
3398+   return why;
3399+}
3400+
3401+static int stbi__process_frame_header(stbi__jpeg *z, int scan)
3402+{
3403+   stbi__context *s = z->s;
3404+   int Lf,p,i,q, h_max=1,v_max=1,c;
3405+   Lf = stbi__get16be(s);         if (Lf < 11) return stbi__err("bad SOF len","Corrupt JPEG"); /* JPEG */
3406+   p  = stbi__get8(s);            if (p != 8) return stbi__err("only 8-bit","JPEG format not supported: 8-bit only"); /* JPEG baseline */
3407+   s->img_y = stbi__get16be(s);   if (s->img_y == 0) return stbi__err("no header height", "JPEG format not supported: delayed height"); /* Legal, but we don't handle it--but neither does IJG */
3408+   s->img_x = stbi__get16be(s);   if (s->img_x == 0) return stbi__err("0 width","Corrupt JPEG"); /* JPEG requires */
3409+   if (s->img_y > STBI_MAX_DIMENSIONS) return stbi__err("too large","Very large image (corrupt?)");
3410+   if (s->img_x > STBI_MAX_DIMENSIONS) return stbi__err("too large","Very large image (corrupt?)");
3411+   c = stbi__get8(s);
3412+   if (c != 3 && c != 1 && c != 4) return stbi__err("bad component count","Corrupt JPEG");
3413+   s->img_n = c;
3414+   for (i=0; i < c; ++i) {
3415+      z->img_comp[i].data = NULL;
3416+      z->img_comp[i].linebuf = NULL;
3417+   }
3418+
3419+   if (Lf != 8+3*s->img_n) return stbi__err("bad SOF len","Corrupt JPEG");
3420+
3421+   z->rgb = 0;
3422+   for (i=0; i < s->img_n; ++i) {
3423+      static const unsigned char rgb[3] = { 'R', 'G', 'B' };
3424+      z->img_comp[i].id = stbi__get8(s);
3425+      if (s->img_n == 3 && z->img_comp[i].id == rgb[i])
3426+         ++z->rgb;
3427+      q = stbi__get8(s);
3428+      z->img_comp[i].h = (q >> 4);  if (!z->img_comp[i].h || z->img_comp[i].h > 4) return stbi__err("bad H","Corrupt JPEG");
3429+      z->img_comp[i].v = q & 15;    if (!z->img_comp[i].v || z->img_comp[i].v > 4) return stbi__err("bad V","Corrupt JPEG");
3430+      z->img_comp[i].tq = stbi__get8(s);  if (z->img_comp[i].tq > 3) return stbi__err("bad TQ","Corrupt JPEG");
3431+   }
3432+
3433+   if (scan != STBI__SCAN_load) return 1;
3434+
3435+   if (!stbi__mad3sizes_valid(s->img_x, s->img_y, s->img_n, 0)) return stbi__err("too large", "Image too large to decode");
3436+
3437+   for (i=0; i < s->img_n; ++i) {
3438+      if (z->img_comp[i].h > h_max) h_max = z->img_comp[i].h;
3439+      if (z->img_comp[i].v > v_max) v_max = z->img_comp[i].v;
3440+   }
3441+
3442+   /*
3443+    * check that plane subsampling factors are integer ratios; our resamplers can't deal with fractional ratios
3444+    * and I've never seen a non-corrupted JPEG file actually use them
3445+    */
3446+   for (i=0; i < s->img_n; ++i) {
3447+      if (h_max % z->img_comp[i].h != 0) return stbi__err("bad H","Corrupt JPEG");
3448+      if (v_max % z->img_comp[i].v != 0) return stbi__err("bad V","Corrupt JPEG");
3449+   }
3450+
3451+   /* compute interleaved mcu info */
3452+   z->img_h_max = h_max;
3453+   z->img_v_max = v_max;
3454+   z->img_mcu_w = h_max * 8;
3455+   z->img_mcu_h = v_max * 8;
3456+   /* these sizes can't be more than 17 bits */
3457+   z->img_mcu_x = (s->img_x + z->img_mcu_w-1) / z->img_mcu_w;
3458+   z->img_mcu_y = (s->img_y + z->img_mcu_h-1) / z->img_mcu_h;
3459+
3460+   for (i=0; i < s->img_n; ++i) {
3461+      /* number of effective pixels (e.g. for non-interleaved MCU) */
3462+      z->img_comp[i].x = (s->img_x * z->img_comp[i].h + h_max-1) / h_max;
3463+      z->img_comp[i].y = (s->img_y * z->img_comp[i].v + v_max-1) / v_max;
3464+      /*
3465+       * to simplify generation, we'll allocate enough memory to decode
3466+       * the bogus oversized data from using interleaved MCUs and their
3467+       * big blocks (e.g. a 16x16 iMCU on an image of width 33); we won't
3468+       * discard the extra data until colorspace conversion
3469+       *
3470+       * img_mcu_x, img_mcu_y: <=17 bits; comp[i].h and .v are <=4 (checked earlier)
3471+       * so these muls can't overflow with 32-bit ints (which we require)
3472+       */
3473+      z->img_comp[i].w2 = z->img_mcu_x * z->img_comp[i].h * 8;
3474+      z->img_comp[i].h2 = z->img_mcu_y * z->img_comp[i].v * 8;
3475+      z->img_comp[i].coeff = 0;
3476+      z->img_comp[i].raw_coeff = 0;
3477+      z->img_comp[i].linebuf = NULL;
3478+      z->img_comp[i].raw_data = stbi__malloc_mad2(z->img_comp[i].w2, z->img_comp[i].h2, 15);
3479+      if (z->img_comp[i].raw_data == NULL)
3480+         return stbi__free_jpeg_components(z, i+1, stbi__err("outofmem", "Out of memory"));
3481+      /* align blocks for idct using mmx/sse */
3482+      z->img_comp[i].data = (stbi_uc*) (((size_t) z->img_comp[i].raw_data + 15) & ~15);
3483+      if (z->progressive) {
3484+         /* w2, h2 are multiples of 8 (see above) */
3485+         z->img_comp[i].coeff_w = z->img_comp[i].w2 / 8;
3486+         z->img_comp[i].coeff_h = z->img_comp[i].h2 / 8;
3487+         z->img_comp[i].raw_coeff = stbi__malloc_mad3(z->img_comp[i].w2, z->img_comp[i].h2, sizeof(short), 15);
3488+         if (z->img_comp[i].raw_coeff == NULL)
3489+            return stbi__free_jpeg_components(z, i+1, stbi__err("outofmem", "Out of memory"));
3490+         z->img_comp[i].coeff = (short*) (((size_t) z->img_comp[i].raw_coeff + 15) & ~15);
3491+      }
3492+   }
3493+
3494+   return 1;
3495+}
3496+
3497+/* use comparisons since in some cases we handle more than one case (e.g. SOF) */
3498+#define stbi__DNL(x)         ((x) == 0xdc)
3499+#define stbi__SOI(x)         ((x) == 0xd8)
3500+#define stbi__EOI(x)         ((x) == 0xd9)
3501+#define stbi__SOF(x)         ((x) == 0xc0 || (x) == 0xc1 || (x) == 0xc2)
3502+#define stbi__SOS(x)         ((x) == 0xda)
3503+
3504+#define stbi__SOF_progressive(x)   ((x) == 0xc2)
3505+
3506+static int stbi__decode_jpeg_header(stbi__jpeg *z, int scan)
3507+{
3508+   int m;
3509+   z->jfif = 0;
3510+   z->app14_color_transform = -1; /* valid values are 0,1,2 */
3511+   z->marker = STBI__MARKER_none; /* initialize cached marker to empty */
3512+   m = stbi__get_marker(z);
3513+   if (!stbi__SOI(m)) return stbi__err("no SOI","Corrupt JPEG");
3514+   if (scan == STBI__SCAN_type) return 1;
3515+   m = stbi__get_marker(z);
3516+   while (!stbi__SOF(m)) {
3517+      if (!stbi__process_marker(z,m)) return 0;
3518+      m = stbi__get_marker(z);
3519+      while (m == STBI__MARKER_none) {
3520+         /* some files have extra padding after their blocks, so ok, we'll scan */
3521+         if (stbi__at_eof(z->s)) return stbi__err("no SOF", "Corrupt JPEG");
3522+         m = stbi__get_marker(z);
3523+      }
3524+   }
3525+   z->progressive = stbi__SOF_progressive(m);
3526+   if (!stbi__process_frame_header(z, scan)) return 0;
3527+   return 1;
3528+}
3529+
3530+static stbi_uc stbi__skip_jpeg_junk_at_end(stbi__jpeg *j)
3531+{
3532+   /*
3533+    * some JPEGs have junk at end, skip over it but if we find what looks
3534+    * like a valid marker, resume there
3535+    */
3536+   while (!stbi__at_eof(j->s)) {
3537+      stbi_uc x = stbi__get8(j->s);
3538+      while (x == 0xff) { /* might be a marker */
3539+         if (stbi__at_eof(j->s)) return STBI__MARKER_none;
3540+         x = stbi__get8(j->s);
3541+         if (x != 0x00 && x != 0xff) {
3542+            /*
3543+             * not a stuffed zero or lead-in to another marker, looks
3544+             * like an actual marker, return it
3545+             */
3546+            return x;
3547+         }
3548+         /*
3549+          * stuffed zero has x=0 now which ends the loop, meaning we go
3550+          * back to regular scan loop.
3551+          * repeated 0xff keeps trying to read the next byte of the marker.
3552+          */
3553+      }
3554+   }
3555+   return STBI__MARKER_none;
3556+}
3557+
3558+/* decode image to YCbCr format */
3559+static int stbi__decode_jpeg_image(stbi__jpeg *j)
3560+{
3561+   int m;
3562+   for (m = 0; m < 4; m++) {
3563+      j->img_comp[m].raw_data = NULL;
3564+      j->img_comp[m].raw_coeff = NULL;
3565+   }
3566+   j->restart_interval = 0;
3567+   if (!stbi__decode_jpeg_header(j, STBI__SCAN_load)) return 0;
3568+   m = stbi__get_marker(j);
3569+   while (!stbi__EOI(m)) {
3570+      if (stbi__SOS(m)) {
3571+         if (!stbi__process_scan_header(j)) return 0;
3572+         if (!stbi__parse_entropy_coded_data(j)) return 0;
3573+         if (j->marker == STBI__MARKER_none ) {
3574+         j->marker = stbi__skip_jpeg_junk_at_end(j);
3575+            /* if we reach eof without hitting a marker, stbi__get_marker() below will fail and we'll eventually return 0 */
3576+         }
3577+         m = stbi__get_marker(j);
3578+         if (STBI__RESTART(m))
3579+            m = stbi__get_marker(j);
3580+      } else if (stbi__DNL(m)) {
3581+         int Ld = stbi__get16be(j->s);
3582+         stbi__uint32 NL = stbi__get16be(j->s);
3583+         if (Ld != 4) return stbi__err("bad DNL len", "Corrupt JPEG");
3584+         if (NL != j->s->img_y) return stbi__err("bad DNL height", "Corrupt JPEG");
3585+         m = stbi__get_marker(j);
3586+      } else {
3587+         if (!stbi__process_marker(j, m)) return 1;
3588+         m = stbi__get_marker(j);
3589+      }
3590+   }
3591+   if (j->progressive)
3592+      stbi__jpeg_finish(j);
3593+   return 1;
3594+}
3595+
3596+/* static jfif-centered resampling (across block boundaries) */
3597+
3598+typedef stbi_uc *(*resample_row_func)(stbi_uc *out, stbi_uc *in0, stbi_uc *in1,
3599+                                    int w, int hs);
3600+
3601+#define stbi__div4(x) ((stbi_uc) ((x) >> 2))
3602+
3603+static stbi_uc *resample_row_1(stbi_uc *out, stbi_uc *in_near, stbi_uc *in_far, int w, int hs)
3604+{
3605+   STBI_NOTUSED(out);
3606+   STBI_NOTUSED(in_far);
3607+   STBI_NOTUSED(w);
3608+   STBI_NOTUSED(hs);
3609+   return in_near;
3610+}
3611+
3612+static stbi_uc* stbi__resample_row_v_2(stbi_uc *out, stbi_uc *in_near, stbi_uc *in_far, int w, int hs)
3613+{
3614+   /* need to generate two samples vertically for every one in input */
3615+   int i;
3616+   STBI_NOTUSED(hs);
3617+   for (i=0; i < w; ++i)
3618+      out[i] = stbi__div4(3*in_near[i] + in_far[i] + 2);
3619+   return out;
3620+}
3621+
3622+static stbi_uc*  stbi__resample_row_h_2(stbi_uc *out, stbi_uc *in_near, stbi_uc *in_far, int w, int hs)
3623+{
3624+   /* need to generate two samples horizontally for every one in input */
3625+   int i;
3626+   stbi_uc *input = in_near;
3627+
3628+   if (w == 1) {
3629+      /* if only one sample, can't do any interpolation */
3630+      out[0] = out[1] = input[0];
3631+      return out;
3632+   }
3633+
3634+   out[0] = input[0];
3635+   out[1] = stbi__div4(input[0]*3 + input[1] + 2);
3636+   for (i=1; i < w-1; ++i) {
3637+      int n = 3*input[i]+2;
3638+      out[i*2+0] = stbi__div4(n+input[i-1]);
3639+      out[i*2+1] = stbi__div4(n+input[i+1]);
3640+   }
3641+   out[i*2+0] = stbi__div4(input[w-2]*3 + input[w-1] + 2);
3642+   out[i*2+1] = input[w-1];
3643+
3644+   STBI_NOTUSED(in_far);
3645+   STBI_NOTUSED(hs);
3646+
3647+   return out;
3648+}
3649+
3650+#define stbi__div16(x) ((stbi_uc) ((x) >> 4))
3651+
3652+static stbi_uc *stbi__resample_row_hv_2(stbi_uc *out, stbi_uc *in_near, stbi_uc *in_far, int w, int hs)
3653+{
3654+   /* need to generate 2x2 samples for every one in input */
3655+   int i,t0,t1;
3656+   if (w == 1) {
3657+      out[0] = out[1] = stbi__div4(3*in_near[0] + in_far[0] + 2);
3658+      return out;
3659+   }
3660+
3661+   t1 = 3*in_near[0] + in_far[0];
3662+   out[0] = stbi__div4(t1+2);
3663+   for (i=1; i < w; ++i) {
3664+      t0 = t1;
3665+      t1 = 3*in_near[i]+in_far[i];
3666+      out[i*2-1] = stbi__div16(3*t0 + t1 + 8);
3667+      out[i*2  ] = stbi__div16(3*t1 + t0 + 8);
3668+   }
3669+   out[w*2-1] = stbi__div4(t1+2);
3670+
3671+   STBI_NOTUSED(hs);
3672+
3673+   return out;
3674+}
3675+
3676+#if defined(STBI_SSE2) || defined(STBI_NEON)
3677+static stbi_uc *stbi__resample_row_hv_2_simd(stbi_uc *out, stbi_uc *in_near, stbi_uc *in_far, int w, int hs)
3678+{
3679+   /* need to generate 2x2 samples for every one in input */
3680+   int i=0,t0,t1;
3681+
3682+   if (w == 1) {
3683+      out[0] = out[1] = stbi__div4(3*in_near[0] + in_far[0] + 2);
3684+      return out;
3685+   }
3686+
3687+   t1 = 3*in_near[0] + in_far[0];
3688+   /*
3689+    * process groups of 8 pixels for as long as we can.
3690+    * note we can't handle the last pixel in a row in this loop
3691+    * because we need to handle the filter boundary conditions.
3692+    */
3693+   for (; i < ((w-1) & ~7); i += 8) {
3694+#if defined(STBI_SSE2)
3695+      /*
3696+       * load and perform the vertical filtering pass
3697+       * this uses 3*x + y = 4*x + (y - x)
3698+       */
3699+      __m128i zero  = _mm_setzero_si128();
3700+      __m128i farb  = _mm_loadl_epi64((__m128i *) (in_far + i));
3701+      __m128i nearb = _mm_loadl_epi64((__m128i *) (in_near + i));
3702+      __m128i farw  = _mm_unpacklo_epi8(farb, zero);
3703+      __m128i nearw = _mm_unpacklo_epi8(nearb, zero);
3704+      __m128i diff  = _mm_sub_epi16(farw, nearw);
3705+      __m128i nears = _mm_slli_epi16(nearw, 2);
3706+      __m128i curr  = _mm_add_epi16(nears, diff); /* current row */
3707+
3708+      /*
3709+       * horizontal filter works the same based on shifted vers of current
3710+       * row. "prev" is current row shifted right by 1 pixel; we need to
3711+       * insert the previous pixel value (from t1).
3712+       * "next" is current row shifted left by 1 pixel, with first pixel
3713+       * of next block of 8 pixels added in.
3714+       */
3715+      __m128i prv0 = _mm_slli_si128(curr, 2);
3716+      __m128i nxt0 = _mm_srli_si128(curr, 2);
3717+      __m128i prev = _mm_insert_epi16(prv0, t1, 0);
3718+      __m128i next = _mm_insert_epi16(nxt0, 3*in_near[i+8] + in_far[i+8], 7);
3719+
3720+      /*
3721+       * horizontal filter, polyphase implementation since it's convenient:
3722+       * even pixels = 3*cur + prev = cur*4 + (prev - cur)
3723+       * odd  pixels = 3*cur + next = cur*4 + (next - cur)
3724+       * note the shared term.
3725+       */
3726+      __m128i bias  = _mm_set1_epi16(8);
3727+      __m128i curs = _mm_slli_epi16(curr, 2);
3728+      __m128i prvd = _mm_sub_epi16(prev, curr);
3729+      __m128i nxtd = _mm_sub_epi16(next, curr);
3730+      __m128i curb = _mm_add_epi16(curs, bias);
3731+      __m128i even = _mm_add_epi16(prvd, curb);
3732+      __m128i odd  = _mm_add_epi16(nxtd, curb);
3733+
3734+      /* interleave even and odd pixels, then undo scaling. */
3735+      __m128i int0 = _mm_unpacklo_epi16(even, odd);
3736+      __m128i int1 = _mm_unpackhi_epi16(even, odd);
3737+      __m128i de0  = _mm_srli_epi16(int0, 4);
3738+      __m128i de1  = _mm_srli_epi16(int1, 4);
3739+
3740+      /* pack and write output */
3741+      __m128i outv = _mm_packus_epi16(de0, de1);
3742+      _mm_storeu_si128((__m128i *) (out + i*2), outv);
3743+#elif defined(STBI_NEON)
3744+      /*
3745+       * load and perform the vertical filtering pass
3746+       * this uses 3*x + y = 4*x + (y - x)
3747+       */
3748+      uint8x8_t farb  = vld1_u8(in_far + i);
3749+      uint8x8_t nearb = vld1_u8(in_near + i);
3750+      int16x8_t diff  = vreinterpretq_s16_u16(vsubl_u8(farb, nearb));
3751+      int16x8_t nears = vreinterpretq_s16_u16(vshll_n_u8(nearb, 2));
3752+      int16x8_t curr  = vaddq_s16(nears, diff); /* current row */
3753+
3754+      /*
3755+       * horizontal filter works the same based on shifted vers of current
3756+       * row. "prev" is current row shifted right by 1 pixel; we need to
3757+       * insert the previous pixel value (from t1).
3758+       * "next" is current row shifted left by 1 pixel, with first pixel
3759+       * of next block of 8 pixels added in.
3760+       */
3761+      int16x8_t prv0 = vextq_s16(curr, curr, 7);
3762+      int16x8_t nxt0 = vextq_s16(curr, curr, 1);
3763+      int16x8_t prev = vsetq_lane_s16(t1, prv0, 0);
3764+      int16x8_t next = vsetq_lane_s16(3*in_near[i+8] + in_far[i+8], nxt0, 7);
3765+
3766+      /*
3767+       * horizontal filter, polyphase implementation since it's convenient:
3768+       * even pixels = 3*cur + prev = cur*4 + (prev - cur)
3769+       * odd  pixels = 3*cur + next = cur*4 + (next - cur)
3770+       * note the shared term.
3771+       */
3772+      int16x8_t curs = vshlq_n_s16(curr, 2);
3773+      int16x8_t prvd = vsubq_s16(prev, curr);
3774+      int16x8_t nxtd = vsubq_s16(next, curr);
3775+      int16x8_t even = vaddq_s16(curs, prvd);
3776+      int16x8_t odd  = vaddq_s16(curs, nxtd);
3777+
3778+      /* undo scaling and round, then store with even/odd phases interleaved */
3779+      uint8x8x2_t o;
3780+      o.val[0] = vqrshrun_n_s16(even, 4);
3781+      o.val[1] = vqrshrun_n_s16(odd,  4);
3782+      vst2_u8(out + i*2, o);
3783+#endif
3784+
3785+      /* "previous" value for next iter */
3786+      t1 = 3*in_near[i+7] + in_far[i+7];
3787+   }
3788+
3789+   t0 = t1;
3790+   t1 = 3*in_near[i] + in_far[i];
3791+   out[i*2] = stbi__div16(3*t1 + t0 + 8);
3792+
3793+   for (++i; i < w; ++i) {
3794+      t0 = t1;
3795+      t1 = 3*in_near[i]+in_far[i];
3796+      out[i*2-1] = stbi__div16(3*t0 + t1 + 8);
3797+      out[i*2  ] = stbi__div16(3*t1 + t0 + 8);
3798+   }
3799+   out[w*2-1] = stbi__div4(t1+2);
3800+
3801+   STBI_NOTUSED(hs);
3802+
3803+   return out;
3804+}
3805+#endif
3806+
3807+static stbi_uc *stbi__resample_row_generic(stbi_uc *out, stbi_uc *in_near, stbi_uc *in_far, int w, int hs)
3808+{
3809+   /* resample with nearest-neighbor */
3810+   int i,j;
3811+   STBI_NOTUSED(in_far);
3812+   for (i=0; i < w; ++i)
3813+      for (j=0; j < hs; ++j)
3814+         out[i*hs+j] = in_near[i];
3815+   return out;
3816+}
3817+
3818+/*
3819+ * this is a reduced-precision calculation of YCbCr-to-RGB introduced
3820+ * to make sure the code produces the same results in both SIMD and scalar
3821+ */
3822+#define stbi__float2fixed(x)  (((int) ((x) * 4096.0f + 0.5f)) << 8)
3823+static void stbi__YCbCr_to_RGB_row(stbi_uc *out, const stbi_uc *y, const stbi_uc *pcb, const stbi_uc *pcr, int count, int step)
3824+{
3825+   int i;
3826+   for (i=0; i < count; ++i) {
3827+      int y_fixed = (y[i] << 20) + (1<<19); /* rounding */
3828+      int r,g,b;
3829+      int cr = pcr[i] - 128;
3830+      int cb = pcb[i] - 128;
3831+      r = y_fixed +  cr* stbi__float2fixed(1.40200f);
3832+      g = y_fixed + (cr*-stbi__float2fixed(0.71414f)) + ((cb*-stbi__float2fixed(0.34414f)) & 0xffff0000);
3833+      b = y_fixed                                     +   cb* stbi__float2fixed(1.77200f);
3834+      r >>= 20;
3835+      g >>= 20;
3836+      b >>= 20;
3837+      if ((unsigned) r > 255) { if (r < 0) r = 0; else r = 255; }
3838+      if ((unsigned) g > 255) { if (g < 0) g = 0; else g = 255; }
3839+      if ((unsigned) b > 255) { if (b < 0) b = 0; else b = 255; }
3840+      out[0] = (stbi_uc)r;
3841+      out[1] = (stbi_uc)g;
3842+      out[2] = (stbi_uc)b;
3843+      out[3] = 255;
3844+      out += step;
3845+   }
3846+}
3847+
3848+#if defined(STBI_SSE2) || defined(STBI_NEON)
3849+static void stbi__YCbCr_to_RGB_simd(stbi_uc *out, stbi_uc const *y, stbi_uc const *pcb, stbi_uc const *pcr, int count, int step)
3850+{
3851+   int i = 0;
3852+
3853+#ifdef STBI_SSE2
3854+   /*
3855+    * step == 3 is pretty ugly on the final interleave, and i'm not convinced
3856+    * it's useful in practice (you wouldn't use it for textures, for example).
3857+    * so just accelerate step == 4 case.
3858+    */
3859+   if (step == 4) {
3860+      /* this is a fairly straightforward implementation and not super-optimized. */
3861+      __m128i signflip  = _mm_set1_epi8(-0x80);
3862+      __m128i cr_const0 = _mm_set1_epi16(   (short) ( 1.40200f*4096.0f+0.5f));
3863+      __m128i cr_const1 = _mm_set1_epi16( - (short) ( 0.71414f*4096.0f+0.5f));
3864+      __m128i cb_const0 = _mm_set1_epi16( - (short) ( 0.34414f*4096.0f+0.5f));
3865+      __m128i cb_const1 = _mm_set1_epi16(   (short) ( 1.77200f*4096.0f+0.5f));
3866+      __m128i y_bias = _mm_set1_epi8((char) (unsigned char) 128);
3867+      __m128i xw = _mm_set1_epi16(255); /* alpha channel */
3868+
3869+      for (; i+7 < count; i += 8) {
3870+         /* load */
3871+         __m128i y_bytes = _mm_loadl_epi64((__m128i *) (y+i));
3872+         __m128i cr_bytes = _mm_loadl_epi64((__m128i *) (pcr+i));
3873+         __m128i cb_bytes = _mm_loadl_epi64((__m128i *) (pcb+i));
3874+         __m128i cr_biased = _mm_xor_si128(cr_bytes, signflip); /* -128 */
3875+         __m128i cb_biased = _mm_xor_si128(cb_bytes, signflip); /* -128 */
3876+
3877+         /* unpack to short (and left-shift cr, cb by 8) */
3878+         __m128i yw  = _mm_unpacklo_epi8(y_bias, y_bytes);
3879+         __m128i crw = _mm_unpacklo_epi8(_mm_setzero_si128(), cr_biased);
3880+         __m128i cbw = _mm_unpacklo_epi8(_mm_setzero_si128(), cb_biased);
3881+
3882+         /* color transform */
3883+         __m128i yws = _mm_srli_epi16(yw, 4);
3884+         __m128i cr0 = _mm_mulhi_epi16(cr_const0, crw);
3885+         __m128i cb0 = _mm_mulhi_epi16(cb_const0, cbw);
3886+         __m128i cb1 = _mm_mulhi_epi16(cbw, cb_const1);
3887+         __m128i cr1 = _mm_mulhi_epi16(crw, cr_const1);
3888+         __m128i rws = _mm_add_epi16(cr0, yws);
3889+         __m128i gwt = _mm_add_epi16(cb0, yws);
3890+         __m128i bws = _mm_add_epi16(yws, cb1);
3891+         __m128i gws = _mm_add_epi16(gwt, cr1);
3892+
3893+         /* descale */
3894+         __m128i rw = _mm_srai_epi16(rws, 4);
3895+         __m128i bw = _mm_srai_epi16(bws, 4);
3896+         __m128i gw = _mm_srai_epi16(gws, 4);
3897+
3898+         /* back to byte, set up for transpose */
3899+         __m128i brb = _mm_packus_epi16(rw, bw);
3900+         __m128i gxb = _mm_packus_epi16(gw, xw);
3901+
3902+         /* transpose to interleave channels */
3903+         __m128i t0 = _mm_unpacklo_epi8(brb, gxb);
3904+         __m128i t1 = _mm_unpackhi_epi8(brb, gxb);
3905+         __m128i o0 = _mm_unpacklo_epi16(t0, t1);
3906+         __m128i o1 = _mm_unpackhi_epi16(t0, t1);
3907+
3908+         /* store */
3909+         _mm_storeu_si128((__m128i *) (out + 0), o0);
3910+         _mm_storeu_si128((__m128i *) (out + 16), o1);
3911+         out += 32;
3912+      }
3913+   }
3914+#endif
3915+
3916+#ifdef STBI_NEON
3917+   /* in this version, step=3 support would be easy to add. but is there demand? */
3918+   if (step == 4) {
3919+      /* this is a fairly straightforward implementation and not super-optimized. */
3920+      uint8x8_t signflip = vdup_n_u8(0x80);
3921+      int16x8_t cr_const0 = vdupq_n_s16(   (short) ( 1.40200f*4096.0f+0.5f));
3922+      int16x8_t cr_const1 = vdupq_n_s16( - (short) ( 0.71414f*4096.0f+0.5f));
3923+      int16x8_t cb_const0 = vdupq_n_s16( - (short) ( 0.34414f*4096.0f+0.5f));
3924+      int16x8_t cb_const1 = vdupq_n_s16(   (short) ( 1.77200f*4096.0f+0.5f));
3925+
3926+      for (; i+7 < count; i += 8) {
3927+         /* load */
3928+         uint8x8_t y_bytes  = vld1_u8(y + i);
3929+         uint8x8_t cr_bytes = vld1_u8(pcr + i);
3930+         uint8x8_t cb_bytes = vld1_u8(pcb + i);
3931+         int8x8_t cr_biased = vreinterpret_s8_u8(vsub_u8(cr_bytes, signflip));
3932+         int8x8_t cb_biased = vreinterpret_s8_u8(vsub_u8(cb_bytes, signflip));
3933+
3934+         /* expand to s16 */
3935+         int16x8_t yws = vreinterpretq_s16_u16(vshll_n_u8(y_bytes, 4));
3936+         int16x8_t crw = vshll_n_s8(cr_biased, 7);
3937+         int16x8_t cbw = vshll_n_s8(cb_biased, 7);
3938+
3939+         /* color transform */
3940+         int16x8_t cr0 = vqdmulhq_s16(crw, cr_const0);
3941+         int16x8_t cb0 = vqdmulhq_s16(cbw, cb_const0);
3942+         int16x8_t cr1 = vqdmulhq_s16(crw, cr_const1);
3943+         int16x8_t cb1 = vqdmulhq_s16(cbw, cb_const1);
3944+         int16x8_t rws = vaddq_s16(yws, cr0);
3945+         int16x8_t gws = vaddq_s16(vaddq_s16(yws, cb0), cr1);
3946+         int16x8_t bws = vaddq_s16(yws, cb1);
3947+
3948+         /* undo scaling, round, convert to byte */
3949+         uint8x8x4_t o;
3950+         o.val[0] = vqrshrun_n_s16(rws, 4);
3951+         o.val[1] = vqrshrun_n_s16(gws, 4);
3952+         o.val[2] = vqrshrun_n_s16(bws, 4);
3953+         o.val[3] = vdup_n_u8(255);
3954+
3955+         /* store, interleaving r/g/b/a */
3956+         vst4_u8(out, o);
3957+         out += 8*4;
3958+      }
3959+   }
3960+#endif
3961+
3962+   for (; i < count; ++i) {
3963+      int y_fixed = (y[i] << 20) + (1<<19); /* rounding */
3964+      int r,g,b;
3965+      int cr = pcr[i] - 128;
3966+      int cb = pcb[i] - 128;
3967+      r = y_fixed + cr* stbi__float2fixed(1.40200f);
3968+      g = y_fixed + cr*-stbi__float2fixed(0.71414f) + ((cb*-stbi__float2fixed(0.34414f)) & 0xffff0000);
3969+      b = y_fixed                                   +   cb* stbi__float2fixed(1.77200f);
3970+      r >>= 20;
3971+      g >>= 20;
3972+      b >>= 20;
3973+      if ((unsigned) r > 255) { if (r < 0) r = 0; else r = 255; }
3974+      if ((unsigned) g > 255) { if (g < 0) g = 0; else g = 255; }
3975+      if ((unsigned) b > 255) { if (b < 0) b = 0; else b = 255; }
3976+      out[0] = (stbi_uc)r;
3977+      out[1] = (stbi_uc)g;
3978+      out[2] = (stbi_uc)b;
3979+      out[3] = 255;
3980+      out += step;
3981+   }
3982+}
3983+#endif
3984+
3985+/* set up the kernels */
3986+static void stbi__setup_jpeg(stbi__jpeg *j)
3987+{
3988+   j->idct_block_kernel = stbi__idct_block;
3989+   j->YCbCr_to_RGB_kernel = stbi__YCbCr_to_RGB_row;
3990+   j->resample_row_hv_2_kernel = stbi__resample_row_hv_2;
3991+
3992+#ifdef STBI_SSE2
3993+   if (stbi__sse2_available()) {
3994+      j->idct_block_kernel = stbi__idct_simd;
3995+      j->YCbCr_to_RGB_kernel = stbi__YCbCr_to_RGB_simd;
3996+      j->resample_row_hv_2_kernel = stbi__resample_row_hv_2_simd;
3997+   }
3998+#endif
3999+
4000+#ifdef STBI_NEON
4001+   j->idct_block_kernel = stbi__idct_simd;
4002+   j->YCbCr_to_RGB_kernel = stbi__YCbCr_to_RGB_simd;
4003+   j->resample_row_hv_2_kernel = stbi__resample_row_hv_2_simd;
4004+#endif
4005+}
4006+
4007+/* clean up the temporary component buffers */
4008+static void stbi__cleanup_jpeg(stbi__jpeg *j)
4009+{
4010+   stbi__free_jpeg_components(j, j->s->img_n, 0);
4011+}
4012+
4013+typedef struct
4014+{
4015+   resample_row_func resample;
4016+   stbi_uc *line0,*line1;
4017+   int hs,vs; /* expansion factor in each axis */
4018+   int w_lores; /* horizontal pixels pre-expansion */
4019+   int ystep; /* how far through vertical expansion we are */
4020+   int ypos; /* which pre-expansion row we're on */
4021+} stbi__resample;
4022+
4023+/* fast 0..255 * 0..255 => 0..255 rounded multiplication */
4024+static stbi_uc stbi__blinn_8x8(stbi_uc x, stbi_uc y)
4025+{
4026+   unsigned int t = x*y + 128;
4027+   return (stbi_uc) ((t + (t >>8)) >> 8);
4028+}
4029+
4030+static stbi_uc *load_jpeg_image(stbi__jpeg *z, int *out_x, int *out_y, int *comp, int req_comp)
4031+{
4032+   int n, decode_n, is_rgb;
4033+   z->s->img_n = 0; /* make stbi__cleanup_jpeg safe */
4034+
4035+   /* validate req_comp */
4036+   if (req_comp < 0 || req_comp > 4) return stbi__errpuc("bad req_comp", "Internal error");
4037+
4038+   /* load a jpeg image from whichever source, but leave in YCbCr format */
4039+   if (!stbi__decode_jpeg_image(z)) { stbi__cleanup_jpeg(z); return NULL; }
4040+
4041+   /* determine actual number of components to generate */
4042+   n = req_comp ? req_comp : z->s->img_n >= 3 ? 3 : 1;
4043+
4044+   is_rgb = z->s->img_n == 3 && (z->rgb == 3 || (z->app14_color_transform == 0 && !z->jfif));
4045+
4046+   if (z->s->img_n == 3 && n < 3 && !is_rgb)
4047+      decode_n = 1;
4048+   else
4049+      decode_n = z->s->img_n;
4050+
4051+   /*
4052+    * nothing to do if no components requested; check this now to avoid
4053+    * accessing uninitialized coutput[0] later
4054+    */
4055+   if (decode_n <= 0) { stbi__cleanup_jpeg(z); return NULL; }
4056+
4057+   /* resample and color-convert */
4058+   {
4059+      int k;
4060+      unsigned int i,j;
4061+      stbi_uc *output;
4062+      stbi_uc *coutput[4] = { NULL, NULL, NULL, NULL };
4063+
4064+      stbi__resample res_comp[4];
4065+
4066+      for (k=0; k < decode_n; ++k) {
4067+         stbi__resample *r = &res_comp[k];
4068+
4069+         /*
4070+          * allocate line buffer big enough for upsampling off the edges
4071+          * with upsample factor of 4
4072+          */
4073+         z->img_comp[k].linebuf = (stbi_uc *) stbi__malloc(z->s->img_x + 3);
4074+         if (!z->img_comp[k].linebuf) { stbi__cleanup_jpeg(z); return stbi__errpuc("outofmem", "Out of memory"); }
4075+
4076+         r->hs      = z->img_h_max / z->img_comp[k].h;
4077+         r->vs      = z->img_v_max / z->img_comp[k].v;
4078+         r->ystep   = r->vs >> 1;
4079+         r->w_lores = (z->s->img_x + r->hs-1) / r->hs;
4080+         r->ypos    = 0;
4081+         r->line0   = r->line1 = z->img_comp[k].data;
4082+
4083+         if      (r->hs == 1 && r->vs == 1) r->resample = resample_row_1;
4084+         else if (r->hs == 1 && r->vs == 2) r->resample = stbi__resample_row_v_2;
4085+         else if (r->hs == 2 && r->vs == 1) r->resample = stbi__resample_row_h_2;
4086+         else if (r->hs == 2 && r->vs == 2) r->resample = z->resample_row_hv_2_kernel;
4087+         else                               r->resample = stbi__resample_row_generic;
4088+      }
4089+
4090+      /* can't error after this so, this is safe */
4091+      output = (stbi_uc *) stbi__malloc_mad3(n, z->s->img_x, z->s->img_y, 1);
4092+      if (!output) { stbi__cleanup_jpeg(z); return stbi__errpuc("outofmem", "Out of memory"); }
4093+
4094+      /* now go ahead and resample */
4095+      for (j=0; j < z->s->img_y; ++j) {
4096+         stbi_uc *out = output + n * z->s->img_x * j;
4097+         for (k=0; k < decode_n; ++k) {
4098+            stbi__resample *r = &res_comp[k];
4099+            int y_bot = r->ystep >= (r->vs >> 1);
4100+            coutput[k] = r->resample(z->img_comp[k].linebuf,
4101+                                     y_bot ? r->line1 : r->line0,
4102+                                     y_bot ? r->line0 : r->line1,
4103+                                     r->w_lores, r->hs);
4104+            if (++r->ystep >= r->vs) {
4105+               r->ystep = 0;
4106+               r->line0 = r->line1;
4107+               if (++r->ypos < z->img_comp[k].y)
4108+                  r->line1 += z->img_comp[k].w2;
4109+            }
4110+         }
4111+         if (n >= 3) {
4112+            stbi_uc *y = coutput[0];
4113+            if (z->s->img_n == 3) {
4114+               if (is_rgb) {
4115+                  for (i=0; i < z->s->img_x; ++i) {
4116+                     out[0] = y[i];
4117+                     out[1] = coutput[1][i];
4118+                     out[2] = coutput[2][i];
4119+                     out[3] = 255;
4120+                     out += n;
4121+                  }
4122+               } else {
4123+                  z->YCbCr_to_RGB_kernel(out, y, coutput[1], coutput[2], z->s->img_x, n);
4124+               }
4125+            } else if (z->s->img_n == 4) {
4126+               if (z->app14_color_transform == 0) { /* CMYK */
4127+                  for (i=0; i < z->s->img_x; ++i) {
4128+                     stbi_uc m = coutput[3][i];
4129+                     out[0] = stbi__blinn_8x8(coutput[0][i], m);
4130+                     out[1] = stbi__blinn_8x8(coutput[1][i], m);
4131+                     out[2] = stbi__blinn_8x8(coutput[2][i], m);
4132+                     out[3] = 255;
4133+                     out += n;
4134+                  }
4135+               } else if (z->app14_color_transform == 2) { /* YCCK */
4136+                  z->YCbCr_to_RGB_kernel(out, y, coutput[1], coutput[2], z->s->img_x, n);
4137+                  for (i=0; i < z->s->img_x; ++i) {
4138+                     stbi_uc m = coutput[3][i];
4139+                     out[0] = stbi__blinn_8x8(255 - out[0], m);
4140+                     out[1] = stbi__blinn_8x8(255 - out[1], m);
4141+                     out[2] = stbi__blinn_8x8(255 - out[2], m);
4142+                     out += n;
4143+                  }
4144+               } else { /* YCbCr + alpha?  Ignore the fourth channel for now */
4145+                  z->YCbCr_to_RGB_kernel(out, y, coutput[1], coutput[2], z->s->img_x, n);
4146+               }
4147+            } else
4148+               for (i=0; i < z->s->img_x; ++i) {
4149+                  out[0] = out[1] = out[2] = y[i];
4150+                  out[3] = 255; /* not used if n==3 */
4151+                  out += n;
4152+               }
4153+         } else {
4154+            if (is_rgb) {
4155+               if (n == 1)
4156+                  for (i=0; i < z->s->img_x; ++i)
4157+                     *out++ = stbi__compute_y(coutput[0][i], coutput[1][i], coutput[2][i]);
4158+               else {
4159+                  for (i=0; i < z->s->img_x; ++i, out += 2) {
4160+                     out[0] = stbi__compute_y(coutput[0][i], coutput[1][i], coutput[2][i]);
4161+                     out[1] = 255;
4162+                  }
4163+               }
4164+            } else if (z->s->img_n == 4 && z->app14_color_transform == 0) {
4165+               for (i=0; i < z->s->img_x; ++i) {
4166+                  stbi_uc m = coutput[3][i];
4167+                  stbi_uc r = stbi__blinn_8x8(coutput[0][i], m);
4168+                  stbi_uc g = stbi__blinn_8x8(coutput[1][i], m);
4169+                  stbi_uc b = stbi__blinn_8x8(coutput[2][i], m);
4170+                  out[0] = stbi__compute_y(r, g, b);
4171+                  out[1] = 255;
4172+                  out += n;
4173+               }
4174+            } else if (z->s->img_n == 4 && z->app14_color_transform == 2) {
4175+               for (i=0; i < z->s->img_x; ++i) {
4176+                  out[0] = stbi__blinn_8x8(255 - coutput[0][i], coutput[3][i]);
4177+                  out[1] = 255;
4178+                  out += n;
4179+               }
4180+            } else {
4181+               stbi_uc *y = coutput[0];
4182+               if (n == 1)
4183+                  for (i=0; i < z->s->img_x; ++i) out[i] = y[i];
4184+               else
4185+                  for (i=0; i < z->s->img_x; ++i) { *out++ = y[i]; *out++ = 255; }
4186+            }
4187+         }
4188+      }
4189+      stbi__cleanup_jpeg(z);
4190+      *out_x = z->s->img_x;
4191+      *out_y = z->s->img_y;
4192+      if (comp) *comp = z->s->img_n >= 3 ? 3 : 1; /* report original components, not output */
4193+      return output;
4194+   }
4195+}
4196+
4197+static void *stbi__jpeg_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri)
4198+{
4199+   unsigned char* result;
4200+   stbi__jpeg* j = (stbi__jpeg*) stbi__malloc(sizeof(stbi__jpeg));
4201+   if (!j) return stbi__errpuc("outofmem", "Out of memory");
4202+   memset(j, 0, sizeof(stbi__jpeg));
4203+   STBI_NOTUSED(ri);
4204+   j->s = s;
4205+   stbi__setup_jpeg(j);
4206+   result = load_jpeg_image(j, x,y,comp,req_comp);
4207+   STBI_FREE(j);
4208+   return result;
4209+}
4210+
4211+static int stbi__jpeg_test(stbi__context *s)
4212+{
4213+   int r;
4214+   stbi__jpeg* j = (stbi__jpeg*)stbi__malloc(sizeof(stbi__jpeg));
4215+   if (!j) return stbi__err("outofmem", "Out of memory");
4216+   memset(j, 0, sizeof(stbi__jpeg));
4217+   j->s = s;
4218+   stbi__setup_jpeg(j);
4219+   r = stbi__decode_jpeg_header(j, STBI__SCAN_type);
4220+   stbi__rewind(s);
4221+   STBI_FREE(j);
4222+   return r;
4223+}
4224+
4225+static int stbi__jpeg_info_raw(stbi__jpeg *j, int *x, int *y, int *comp)
4226+{
4227+   if (!stbi__decode_jpeg_header(j, STBI__SCAN_header)) {
4228+      stbi__rewind( j->s );
4229+      return 0;
4230+   }
4231+   if (x) *x = j->s->img_x;
4232+   if (y) *y = j->s->img_y;
4233+   if (comp) *comp = j->s->img_n >= 3 ? 3 : 1;
4234+   return 1;
4235+}
4236+
4237+static int stbi__jpeg_info(stbi__context *s, int *x, int *y, int *comp)
4238+{
4239+   int result;
4240+   stbi__jpeg* j = (stbi__jpeg*) (stbi__malloc(sizeof(stbi__jpeg)));
4241+   if (!j) return stbi__err("outofmem", "Out of memory");
4242+   memset(j, 0, sizeof(stbi__jpeg));
4243+   j->s = s;
4244+   result = stbi__jpeg_info_raw(j, x, y, comp);
4245+   STBI_FREE(j);
4246+   return result;
4247+}
4248+#endif
4249+
4250+/*
4251+ * public domain zlib decode    v0.2  Sean Barrett 2006-11-18
4252+ *    simple implementation
4253+ *      - all input must be provided in an upfront buffer
4254+ *      - all output is written to a single output buffer (can malloc/realloc)
4255+ *    performance
4256+ *      - fast huffman
4257+ */
4258+
4259+#ifndef STBI_NO_ZLIB
4260+
4261+/* fast-way is faster to check than jpeg huffman, but slow way is slower */
4262+#define STBI__ZFAST_BITS  9 /* accelerate all cases in default tables */
4263+#define STBI__ZFAST_MASK  ((1 << STBI__ZFAST_BITS) - 1)
4264+#define STBI__ZNSYMS 288 /* number of symbols in literal/length alphabet */
4265+
4266+/*
4267+ * zlib-style huffman encoding
4268+ * (jpegs packs from left, zlib from right, so can't share code)
4269+ */
4270+typedef struct
4271+{
4272+   stbi__uint16 fast[1 << STBI__ZFAST_BITS];
4273+   stbi__uint16 firstcode[16];
4274+   int maxcode[17];
4275+   stbi__uint16 firstsymbol[16];
4276+   stbi_uc  size[STBI__ZNSYMS];
4277+   stbi__uint16 value[STBI__ZNSYMS];
4278+} stbi__zhuffman;
4279+
4280+stbi_inline static int stbi__bitreverse16(int n)
4281+{
4282+  n = ((n & 0xAAAA) >>  1) | ((n & 0x5555) << 1);
4283+  n = ((n & 0xCCCC) >>  2) | ((n & 0x3333) << 2);
4284+  n = ((n & 0xF0F0) >>  4) | ((n & 0x0F0F) << 4);
4285+  n = ((n & 0xFF00) >>  8) | ((n & 0x00FF) << 8);
4286+  return n;
4287+}
4288+
4289+stbi_inline static int stbi__bit_reverse(int v, int bits)
4290+{
4291+   STBI_ASSERT(bits <= 16);
4292+   /*
4293+    * to bit reverse n bits, reverse 16 and shift
4294+    * e.g. 11 bits, bit reverse and shift away 5
4295+    */
4296+   return stbi__bitreverse16(v) >> (16-bits);
4297+}
4298+
4299+static int stbi__zbuild_huffman(stbi__zhuffman *z, const stbi_uc *sizelist, int num)
4300+{
4301+   int i,k=0;
4302+   int code, next_code[16], sizes[17];
4303+
4304+   /* DEFLATE spec for generating codes */
4305+   memset(sizes, 0, sizeof(sizes));
4306+   memset(z->fast, 0, sizeof(z->fast));
4307+   for (i=0; i < num; ++i)
4308+      ++sizes[sizelist[i]];
4309+   sizes[0] = 0;
4310+   for (i=1; i < 16; ++i)
4311+      if (sizes[i] > (1 << i))
4312+         return stbi__err("bad sizes", "Corrupt PNG");
4313+   code = 0;
4314+   for (i=1; i < 16; ++i) {
4315+      next_code[i] = code;
4316+      z->firstcode[i] = (stbi__uint16) code;
4317+      z->firstsymbol[i] = (stbi__uint16) k;
4318+      code = (code + sizes[i]);
4319+      if (sizes[i])
4320+         if (code-1 >= (1 << i)) return stbi__err("bad codelengths","Corrupt PNG");
4321+      z->maxcode[i] = code << (16-i); /* preshift for inner loop */
4322+      code <<= 1;
4323+      k += sizes[i];
4324+   }
4325+   z->maxcode[16] = 0x10000; /* sentinel */
4326+   for (i=0; i < num; ++i) {
4327+      int s = sizelist[i];
4328+      if (s) {
4329+         int c = next_code[s] - z->firstcode[s] + z->firstsymbol[s];
4330+         stbi__uint16 fastv = (stbi__uint16) ((s << 9) | i);
4331+         z->size [c] = (stbi_uc     ) s;
4332+         z->value[c] = (stbi__uint16) i;
4333+         if (s <= STBI__ZFAST_BITS) {
4334+            int j = stbi__bit_reverse(next_code[s],s);
4335+            while (j < (1 << STBI__ZFAST_BITS)) {
4336+               z->fast[j] = fastv;
4337+               j += (1 << s);
4338+            }
4339+         }
4340+         ++next_code[s];
4341+      }
4342+   }
4343+   return 1;
4344+}
4345+
4346+/*
4347+ * zlib-from-memory implementation for PNG reading
4348+ *    because PNG allows splitting the zlib stream arbitrarily,
4349+ *    and it's annoying structurally to have PNG call ZLIB call PNG,
4350+ *    we require PNG read all the IDATs and combine them into a single
4351+ *    memory buffer
4352+ */
4353+
4354+typedef struct
4355+{
4356+   stbi_uc *zbuffer, *zbuffer_end;
4357+   int num_bits;
4358+   int hit_zeof_once;
4359+   stbi__uint32 code_buffer;
4360+
4361+   char *zout;
4362+   char *zout_start;
4363+   char *zout_end;
4364+   int   z_expandable;
4365+
4366+   stbi__zhuffman z_length, z_distance;
4367+} stbi__zbuf;
4368+
4369+stbi_inline static int stbi__zeof(stbi__zbuf *z)
4370+{
4371+   return (z->zbuffer >= z->zbuffer_end);
4372+}
4373+
4374+stbi_inline static stbi_uc stbi__zget8(stbi__zbuf *z)
4375+{
4376+   return stbi__zeof(z) ? 0 : *z->zbuffer++;
4377+}
4378+
4379+static void stbi__fill_bits(stbi__zbuf *z)
4380+{
4381+   do {
4382+      if (z->code_buffer >= (1U << z->num_bits)) {
4383+        z->zbuffer = z->zbuffer_end;  /* treat this as EOF so we fail. */
4384+        return;
4385+      }
4386+      z->code_buffer |= (unsigned int) stbi__zget8(z) << z->num_bits;
4387+      z->num_bits += 8;
4388+   } while (z->num_bits <= 24);
4389+}
4390+
4391+stbi_inline static unsigned int stbi__zreceive(stbi__zbuf *z, int n)
4392+{
4393+   unsigned int k;
4394+   if (z->num_bits < n) stbi__fill_bits(z);
4395+   k = z->code_buffer & ((1 << n) - 1);
4396+   z->code_buffer >>= n;
4397+   z->num_bits -= n;
4398+   return k;
4399+}
4400+
4401+static int stbi__zhuffman_decode_slowpath(stbi__zbuf *a, stbi__zhuffman *z)
4402+{
4403+   int b,s,k;
4404+   /*
4405+    * not resolved by fast table, so compute it the slow way
4406+    * use jpeg approach, which requires MSbits at top
4407+    */
4408+   k = stbi__bit_reverse(a->code_buffer, 16);
4409+   for (s=STBI__ZFAST_BITS+1; ; ++s)
4410+      if (k < z->maxcode[s])
4411+         break;
4412+   if (s >= 16) return -1; /* invalid code! */
4413+   /* code size is s, so: */
4414+   b = (k >> (16-s)) - z->firstcode[s] + z->firstsymbol[s];
4415+   if (b >= STBI__ZNSYMS) return -1; /* some data was corrupt somewhere! */
4416+   if (z->size[b] != s) return -1; /* was originally an assert, but report failure instead. */
4417+   a->code_buffer >>= s;
4418+   a->num_bits -= s;
4419+   return z->value[b];
4420+}
4421+
4422+stbi_inline static int stbi__zhuffman_decode(stbi__zbuf *a, stbi__zhuffman *z)
4423+{
4424+   int b,s;
4425+   if (a->num_bits < 16) {
4426+      if (stbi__zeof(a)) {
4427+         if (!a->hit_zeof_once) {
4428+            /*
4429+             * This is the first time we hit eof, insert 16 extra padding btis
4430+             * to allow us to keep going; if we actually consume any of them
4431+             * though, that is invalid data. This is caught later.
4432+             */
4433+            a->hit_zeof_once = 1;
4434+            a->num_bits += 16; /* add 16 implicit zero bits */
4435+         } else {
4436+            /*
4437+             * We already inserted our extra 16 padding bits and are again
4438+             * out, this stream is actually prematurely terminated.
4439+             */
4440+            return -1;
4441+         }
4442+      } else {
4443+         stbi__fill_bits(a);
4444+      }
4445+   }
4446+   b = z->fast[a->code_buffer & STBI__ZFAST_MASK];
4447+   if (b) {
4448+      s = b >> 9;
4449+      a->code_buffer >>= s;
4450+      a->num_bits -= s;
4451+      return b & 511;
4452+   }
4453+   return stbi__zhuffman_decode_slowpath(a, z);
4454+}
4455+
4456+static int stbi__zexpand(stbi__zbuf *z, char *zout, int n) /* need to make room for n bytes */
4457+{
4458+   char *q;
4459+   unsigned int cur, limit, old_limit;
4460+   z->zout = zout;
4461+   if (!z->z_expandable) return stbi__err("output buffer limit","Corrupt PNG");
4462+   cur   = (unsigned int) (z->zout - z->zout_start);
4463+   limit = old_limit = (unsigned) (z->zout_end - z->zout_start);
4464+   if (UINT_MAX - cur < (unsigned) n) return stbi__err("outofmem", "Out of memory");
4465+   while (cur + n > limit) {
4466+      if(limit > UINT_MAX / 2) return stbi__err("outofmem", "Out of memory");
4467+      limit *= 2;
4468+   }
4469+   q = (char *) STBI_REALLOC_SIZED(z->zout_start, old_limit, limit);
4470+   STBI_NOTUSED(old_limit);
4471+   if (q == NULL) return stbi__err("outofmem", "Out of memory");
4472+   z->zout_start = q;
4473+   z->zout       = q + cur;
4474+   z->zout_end   = q + limit;
4475+   return 1;
4476+}
4477+
4478+static const int stbi__zlength_base[31] = {
4479+   3,4,5,6,7,8,9,10,11,13,
4480+   15,17,19,23,27,31,35,43,51,59,
4481+   67,83,99,115,131,163,195,227,258,0,0 };
4482+
4483+static const int stbi__zlength_extra[31]=
4484+{ 0,0,0,0,0,0,0,0,1,1,1,1,2,2,2,2,3,3,3,3,4,4,4,4,5,5,5,5,0,0,0 };
4485+
4486+static const int stbi__zdist_base[32] = { 1,2,3,4,5,7,9,13,17,25,33,49,65,97,129,193,
4487+257,385,513,769,1025,1537,2049,3073,4097,6145,8193,12289,16385,24577,0,0};
4488+
4489+static const int stbi__zdist_extra[32] =
4490+{ 0,0,0,0,1,1,2,2,3,3,4,4,5,5,6,6,7,7,8,8,9,9,10,10,11,11,12,12,13,13};
4491+
4492+static int stbi__parse_huffman_block(stbi__zbuf *a)
4493+{
4494+   char *zout = a->zout;
4495+   for(;;) {
4496+      int z = stbi__zhuffman_decode(a, &a->z_length);
4497+      if (z < 256) {
4498+         if (z < 0) return stbi__err("bad huffman code","Corrupt PNG"); /* error in huffman codes */
4499+         if (zout >= a->zout_end) {
4500+            if (!stbi__zexpand(a, zout, 1)) return 0;
4501+            zout = a->zout;
4502+         }
4503+         *zout++ = (char) z;
4504+      } else {
4505+         stbi_uc *p;
4506+         int len,dist;
4507+         if (z == 256) {
4508+            a->zout = zout;
4509+            if (a->hit_zeof_once && a->num_bits < 16) {
4510+               /*
4511+                * The first time we hit zeof, we inserted 16 extra zero bits into our bit
4512+                * buffer so the decoder can just do its speculative decoding. But if we
4513+                * actually consumed any of those bits (which is the case when num_bits < 16),
4514+                * the stream actually read past the end so it is malformed.
4515+                */
4516+               return stbi__err("unexpected end","Corrupt PNG");
4517+            }
4518+            return 1;
4519+         }
4520+         if (z >= 286) return stbi__err("bad huffman code","Corrupt PNG"); /* per DEFLATE, length codes 286 and 287 must not appear in compressed data */
4521+         z -= 257;
4522+         len = stbi__zlength_base[z];
4523+         if (stbi__zlength_extra[z]) len += stbi__zreceive(a, stbi__zlength_extra[z]);
4524+         z = stbi__zhuffman_decode(a, &a->z_distance);
4525+         if (z < 0 || z >= 30) return stbi__err("bad huffman code","Corrupt PNG"); /* per DEFLATE, distance codes 30 and 31 must not appear in compressed data */
4526+         dist = stbi__zdist_base[z];
4527+         if (stbi__zdist_extra[z]) dist += stbi__zreceive(a, stbi__zdist_extra[z]);
4528+         if (zout - a->zout_start < dist) return stbi__err("bad dist","Corrupt PNG");
4529+         if (len > a->zout_end - zout) {
4530+            if (!stbi__zexpand(a, zout, len)) return 0;
4531+            zout = a->zout;
4532+         }
4533+         p = (stbi_uc *) (zout - dist);
4534+         if (dist == 1) { /* run of one byte; common in images. */
4535+            stbi_uc v = *p;
4536+            if (len) { do *zout++ = v; while (--len); }
4537+         } else {
4538+            if (len) { do *zout++ = *p++; while (--len); }
4539+         }
4540+      }
4541+   }
4542+}
4543+
4544+static int stbi__compute_huffman_codes(stbi__zbuf *a)
4545+{
4546+   static const stbi_uc length_dezigzag[19] = { 16,17,18,0,8,7,9,6,10,5,11,4,12,3,13,2,14,1,15 };
4547+   stbi__zhuffman z_codelength;
4548+   stbi_uc lencodes[286+32+137]; /* padding for maximum single op */
4549+   stbi_uc codelength_sizes[19];
4550+   int i,n;
4551+
4552+   int hlit  = stbi__zreceive(a,5) + 257;
4553+   int hdist = stbi__zreceive(a,5) + 1;
4554+   int hclen = stbi__zreceive(a,4) + 4;
4555+   int ntot  = hlit + hdist;
4556+
4557+   memset(codelength_sizes, 0, sizeof(codelength_sizes));
4558+   for (i=0; i < hclen; ++i) {
4559+      int s = stbi__zreceive(a,3);
4560+      codelength_sizes[length_dezigzag[i]] = (stbi_uc) s;
4561+   }
4562+   if (!stbi__zbuild_huffman(&z_codelength, codelength_sizes, 19)) return 0;
4563+
4564+   n = 0;
4565+   while (n < ntot) {
4566+      int c = stbi__zhuffman_decode(a, &z_codelength);
4567+      if (c < 0 || c >= 19) return stbi__err("bad codelengths", "Corrupt PNG");
4568+      if (c < 16)
4569+         lencodes[n++] = (stbi_uc) c;
4570+      else {
4571+         stbi_uc fill = 0;
4572+         if (c == 16) {
4573+            c = stbi__zreceive(a,2)+3;
4574+            if (n == 0) return stbi__err("bad codelengths", "Corrupt PNG");
4575+            fill = lencodes[n-1];
4576+         } else if (c == 17) {
4577+            c = stbi__zreceive(a,3)+3;
4578+         } else if (c == 18) {
4579+            c = stbi__zreceive(a,7)+11;
4580+         } else {
4581+            return stbi__err("bad codelengths", "Corrupt PNG");
4582+         }
4583+         if (ntot - n < c) return stbi__err("bad codelengths", "Corrupt PNG");
4584+         memset(lencodes+n, fill, c);
4585+         n += c;
4586+      }
4587+   }
4588+   if (n != ntot) return stbi__err("bad codelengths","Corrupt PNG");
4589+   if (!stbi__zbuild_huffman(&a->z_length, lencodes, hlit)) return 0;
4590+   if (!stbi__zbuild_huffman(&a->z_distance, lencodes+hlit, hdist)) return 0;
4591+   return 1;
4592+}
4593+
4594+static int stbi__parse_uncompressed_block(stbi__zbuf *a)
4595+{
4596+   stbi_uc header[4];
4597+   int len,nlen,k;
4598+   if (a->num_bits & 7)
4599+      stbi__zreceive(a, a->num_bits & 7); /* discard */
4600+   /* drain the bit-packed data into header */
4601+   k = 0;
4602+   while (a->num_bits > 0) {
4603+      header[k++] = (stbi_uc) (a->code_buffer & 255); /* suppress MSVC run-time check */
4604+      a->code_buffer >>= 8;
4605+      a->num_bits -= 8;
4606+   }
4607+   if (a->num_bits < 0) return stbi__err("zlib corrupt","Corrupt PNG");
4608+   /* now fill header the normal way */
4609+   while (k < 4)
4610+      header[k++] = stbi__zget8(a);
4611+   len  = header[1] * 256 + header[0];
4612+   nlen = header[3] * 256 + header[2];
4613+   if (nlen != (len ^ 0xffff)) return stbi__err("zlib corrupt","Corrupt PNG");
4614+   if (a->zbuffer + len > a->zbuffer_end) return stbi__err("read past buffer","Corrupt PNG");
4615+   if (a->zout + len > a->zout_end)
4616+      if (!stbi__zexpand(a, a->zout, len)) return 0;
4617+   memcpy(a->zout, a->zbuffer, len);
4618+   a->zbuffer += len;
4619+   a->zout += len;
4620+   return 1;
4621+}
4622+
4623+static int stbi__parse_zlib_header(stbi__zbuf *a)
4624+{
4625+   int cmf   = stbi__zget8(a);
4626+   int cm    = cmf & 15;
4627+   /* int cinfo = cmf >> 4; */
4628+   int flg   = stbi__zget8(a);
4629+   if (stbi__zeof(a)) return stbi__err("bad zlib header","Corrupt PNG"); /* zlib spec */
4630+   if ((cmf*256+flg) % 31 != 0) return stbi__err("bad zlib header","Corrupt PNG"); /* zlib spec */
4631+   if (flg & 32) return stbi__err("no preset dict","Corrupt PNG"); /* preset dictionary not allowed in png */
4632+   if (cm != 8) return stbi__err("bad compression","Corrupt PNG"); /* DEFLATE required for png */
4633+   /* window = 1 << (8 + cinfo)... but who cares, we fully buffer output */
4634+   return 1;
4635+}
4636+
4637+static const stbi_uc stbi__zdefault_length[STBI__ZNSYMS] =
4638+{
4639+   8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8, 8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,
4640+   8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8, 8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,
4641+   8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8, 8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,
4642+   8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8, 8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,
4643+   8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8, 9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,
4644+   9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,9, 9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,
4645+   9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,9, 9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,
4646+   9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,9, 9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,9,
4647+   7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7, 7,7,7,7,7,7,7,7,8,8,8,8,8,8,8,8
4648+};
4649+static const stbi_uc stbi__zdefault_distance[32] =
4650+{
4651+   5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5
4652+};
4653+/*
4654+Init algorithm:
4655+{
4656+   int i;   // use <= to match clearly with spec
4657+   for (i=0; i <= 143; ++i)     stbi__zdefault_length[i]   = 8;
4658+   for (   ; i <= 255; ++i)     stbi__zdefault_length[i]   = 9;
4659+   for (   ; i <= 279; ++i)     stbi__zdefault_length[i]   = 7;
4660+   for (   ; i <= 287; ++i)     stbi__zdefault_length[i]   = 8;
4661+
4662+   for (i=0; i <=  31; ++i)     stbi__zdefault_distance[i] = 5;
4663+}
4664+*/
4665+
4666+static int stbi__parse_zlib(stbi__zbuf *a, int parse_header)
4667+{
4668+   int final, type;
4669+   if (parse_header)
4670+      if (!stbi__parse_zlib_header(a)) return 0;
4671+   a->num_bits = 0;
4672+   a->code_buffer = 0;
4673+   a->hit_zeof_once = 0;
4674+   do {
4675+      final = stbi__zreceive(a,1);
4676+      type = stbi__zreceive(a,2);
4677+      if (type == 0) {
4678+         if (!stbi__parse_uncompressed_block(a)) return 0;
4679+      } else if (type == 3) {
4680+         return 0;
4681+      } else {
4682+         if (type == 1) {
4683+            /* use fixed code lengths */
4684+            if (!stbi__zbuild_huffman(&a->z_length  , stbi__zdefault_length  , STBI__ZNSYMS)) return 0;
4685+            if (!stbi__zbuild_huffman(&a->z_distance, stbi__zdefault_distance,  32)) return 0;
4686+         } else {
4687+            if (!stbi__compute_huffman_codes(a)) return 0;
4688+         }
4689+         if (!stbi__parse_huffman_block(a)) return 0;
4690+      }
4691+   } while (!final);
4692+   return 1;
4693+}
4694+
4695+static int stbi__do_zlib(stbi__zbuf *a, char *obuf, int olen, int exp, int parse_header)
4696+{
4697+   a->zout_start = obuf;
4698+   a->zout       = obuf;
4699+   a->zout_end   = obuf + olen;
4700+   a->z_expandable = exp;
4701+
4702+   return stbi__parse_zlib(a, parse_header);
4703+}
4704+
4705+STBIDEF char *stbi_zlib_decode_malloc_guesssize(const char *buffer, int len, int initial_size, int *outlen)
4706+{
4707+   stbi__zbuf a;
4708+   char *p = (char *) stbi__malloc(initial_size);
4709+   if (p == NULL) return NULL;
4710+   a.zbuffer = (stbi_uc *) buffer;
4711+   a.zbuffer_end = (stbi_uc *) buffer + len;
4712+   if (stbi__do_zlib(&a, p, initial_size, 1, 1)) {
4713+      if (outlen) *outlen = (int) (a.zout - a.zout_start);
4714+      return a.zout_start;
4715+   } else {
4716+      STBI_FREE(a.zout_start);
4717+      return NULL;
4718+   }
4719+}
4720+
4721+STBIDEF char *stbi_zlib_decode_malloc(char const *buffer, int len, int *outlen)
4722+{
4723+   return stbi_zlib_decode_malloc_guesssize(buffer, len, 16384, outlen);
4724+}
4725+
4726+STBIDEF char *stbi_zlib_decode_malloc_guesssize_headerflag(const char *buffer, int len, int initial_size, int *outlen, int parse_header)
4727+{
4728+   stbi__zbuf a;
4729+   char *p = (char *) stbi__malloc(initial_size);
4730+   if (p == NULL) return NULL;
4731+   a.zbuffer = (stbi_uc *) buffer;
4732+   a.zbuffer_end = (stbi_uc *) buffer + len;
4733+   if (stbi__do_zlib(&a, p, initial_size, 1, parse_header)) {
4734+      if (outlen) *outlen = (int) (a.zout - a.zout_start);
4735+      return a.zout_start;
4736+   } else {
4737+      STBI_FREE(a.zout_start);
4738+      return NULL;
4739+   }
4740+}
4741+
4742+STBIDEF int stbi_zlib_decode_buffer(char *obuffer, int olen, char const *ibuffer, int ilen)
4743+{
4744+   stbi__zbuf a;
4745+   a.zbuffer = (stbi_uc *) ibuffer;
4746+   a.zbuffer_end = (stbi_uc *) ibuffer + ilen;
4747+   if (stbi__do_zlib(&a, obuffer, olen, 0, 1))
4748+      return (int) (a.zout - a.zout_start);
4749+   else
4750+      return -1;
4751+}
4752+
4753+STBIDEF char *stbi_zlib_decode_noheader_malloc(char const *buffer, int len, int *outlen)
4754+{
4755+   stbi__zbuf a;
4756+   char *p = (char *) stbi__malloc(16384);
4757+   if (p == NULL) return NULL;
4758+   a.zbuffer = (stbi_uc *) buffer;
4759+   a.zbuffer_end = (stbi_uc *) buffer+len;
4760+   if (stbi__do_zlib(&a, p, 16384, 1, 0)) {
4761+      if (outlen) *outlen = (int) (a.zout - a.zout_start);
4762+      return a.zout_start;
4763+   } else {
4764+      STBI_FREE(a.zout_start);
4765+      return NULL;
4766+   }
4767+}
4768+
4769+STBIDEF int stbi_zlib_decode_noheader_buffer(char *obuffer, int olen, const char *ibuffer, int ilen)
4770+{
4771+   stbi__zbuf a;
4772+   a.zbuffer = (stbi_uc *) ibuffer;
4773+   a.zbuffer_end = (stbi_uc *) ibuffer + ilen;
4774+   if (stbi__do_zlib(&a, obuffer, olen, 0, 0))
4775+      return (int) (a.zout - a.zout_start);
4776+   else
4777+      return -1;
4778+}
4779+#endif
4780+
4781+/*
4782+ * public domain "baseline" PNG decoder   v0.10  Sean Barrett 2006-11-18
4783+ *    simple implementation
4784+ *      - only 8-bit samples
4785+ *      - no CRC checking
4786+ *      - allocates lots of intermediate memory
4787+ *        - avoids problem of streaming data between subsystems
4788+ *        - avoids explicit window management
4789+ *    performance
4790+ *      - uses stb_zlib, a PD zlib implementation with fast huffman decoding
4791+ */
4792+
4793+#ifndef STBI_NO_PNG
4794+typedef struct
4795+{
4796+   stbi__uint32 length;
4797+   stbi__uint32 type;
4798+} stbi__pngchunk;
4799+
4800+static stbi__pngchunk stbi__get_chunk_header(stbi__context *s)
4801+{
4802+   stbi__pngchunk c;
4803+   c.length = stbi__get32be(s);
4804+   c.type   = stbi__get32be(s);
4805+   return c;
4806+}
4807+
4808+static int stbi__check_png_header(stbi__context *s)
4809+{
4810+   static const stbi_uc png_sig[8] = { 137,80,78,71,13,10,26,10 };
4811+   int i;
4812+   for (i=0; i < 8; ++i)
4813+      if (stbi__get8(s) != png_sig[i]) return stbi__err("bad png sig","Not a PNG");
4814+   return 1;
4815+}
4816+
4817+typedef struct
4818+{
4819+   stbi__context *s;
4820+   stbi_uc *idata, *expanded, *out;
4821+   int depth;
4822+} stbi__png;
4823+
4824+
4825+enum {
4826+   STBI__F_none=0,
4827+   STBI__F_sub=1,
4828+   STBI__F_up=2,
4829+   STBI__F_avg=3,
4830+   STBI__F_paeth=4,
4831+   /* synthetic filter used for first scanline to avoid needing a dummy row of 0s */
4832+   STBI__F_avg_first
4833+};
4834+
4835+static stbi_uc first_row_filter[5] =
4836+{
4837+   STBI__F_none,
4838+   STBI__F_sub,
4839+   STBI__F_none,
4840+   STBI__F_avg_first,
4841+   STBI__F_sub /* Paeth with b=c=0 turns out to be equivalent to sub */
4842+};
4843+
4844+static int stbi__paeth(int a, int b, int c)
4845+{
4846+   /*
4847+    * This formulation looks very different from the reference in the PNG spec, but is
4848+    * actually equivalent and has favorable data dependencies and admits straightforward
4849+    * generation of branch-free code, which helps performance significantly.
4850+    */
4851+   int thresh = c*3 - (a + b);
4852+   int lo = a < b ? a : b;
4853+   int hi = a < b ? b : a;
4854+   int t0 = (hi <= thresh) ? lo : c;
4855+   int t1 = (thresh <= lo) ? hi : t0;
4856+   return t1;
4857+}
4858+
4859+static const stbi_uc stbi__depth_scale_table[9] = { 0, 0xff, 0x55, 0, 0x11, 0,0,0, 0x01 };
4860+
4861+/*
4862+ * adds an extra all-255 alpha channel
4863+ * dest == src is legal
4864+ * img_n must be 1 or 3
4865+ */
4866+static void stbi__create_png_alpha_expand8(stbi_uc *dest, stbi_uc *src, stbi__uint32 x, int img_n)
4867+{
4868+   int i;
4869+   /* must process data backwards since we allow dest==src */
4870+   if (img_n == 1) {
4871+      for (i=x-1; i >= 0; --i) {
4872+         dest[i*2+1] = 255;
4873+         dest[i*2+0] = src[i];
4874+      }
4875+   } else {
4876+      STBI_ASSERT(img_n == 3);
4877+      for (i=x-1; i >= 0; --i) {
4878+         dest[i*4+3] = 255;
4879+         dest[i*4+2] = src[i*3+2];
4880+         dest[i*4+1] = src[i*3+1];
4881+         dest[i*4+0] = src[i*3+0];
4882+      }
4883+   }
4884+}
4885+
4886+/* create the png data from post-deflated data */
4887+static int stbi__create_png_image_raw(stbi__png *a, stbi_uc *raw, stbi__uint32 raw_len, int out_n, stbi__uint32 x, stbi__uint32 y, int depth, int color)
4888+{
4889+   int bytes = (depth == 16 ? 2 : 1);
4890+   stbi__context *s = a->s;
4891+   stbi__uint32 i,j,stride = x*out_n*bytes;
4892+   stbi__uint32 img_len, img_width_bytes;
4893+   stbi_uc *filter_buf;
4894+   int all_ok = 1;
4895+   int k;
4896+   int img_n = s->img_n; /* copy it into a local for later */
4897+
4898+   int output_bytes = out_n*bytes;
4899+   int filter_bytes = img_n*bytes;
4900+   int width = x;
4901+
4902+   STBI_ASSERT(out_n == s->img_n || out_n == s->img_n+1);
4903+   a->out = (stbi_uc *) stbi__malloc_mad3(x, y, output_bytes, 0); /* extra bytes to write off the end into */
4904+   if (!a->out) return stbi__err("outofmem", "Out of memory");
4905+
4906+   /*
4907+    * note: error exits here don't need to clean up a->out individually,
4908+    * stbi__do_png always does on error.
4909+    */
4910+   if (!stbi__mad3sizes_valid(img_n, x, depth, 7)) return stbi__err("too large", "Corrupt PNG");
4911+   img_width_bytes = (((img_n * x * depth) + 7) >> 3);
4912+   if (!stbi__mad2sizes_valid(img_width_bytes, y, img_width_bytes)) return stbi__err("too large", "Corrupt PNG");
4913+   img_len = (img_width_bytes + 1) * y;
4914+
4915+   /*
4916+    * we used to check for exact match between raw_len and img_len on non-interlaced PNGs,
4917+    * but issue #276 reported a PNG in the wild that had extra data at the end (all zeros),
4918+    * so just check for raw_len < img_len always.
4919+    */
4920+   if (raw_len < img_len) return stbi__err("not enough pixels","Corrupt PNG");
4921+
4922+   /* Allocate two scan lines worth of filter workspace buffer. */
4923+   filter_buf = (stbi_uc *) stbi__malloc_mad2(img_width_bytes, 2, 0);
4924+   if (!filter_buf) return stbi__err("outofmem", "Out of memory");
4925+
4926+   /* Filtering for low-bit-depth images */
4927+   if (depth < 8) {
4928+      filter_bytes = 1;
4929+      width = img_width_bytes;
4930+   }
4931+
4932+   for (j=0; j < y; ++j) {
4933+      /* cur/prior filter buffers alternate */
4934+      stbi_uc *cur = filter_buf + (j & 1)*img_width_bytes;
4935+      stbi_uc *prior = filter_buf + (~j & 1)*img_width_bytes;
4936+      stbi_uc *dest = a->out + stride*j;
4937+      int nk = width * filter_bytes;
4938+      int filter = *raw++;
4939+
4940+      /* check filter type */
4941+      if (filter > 4) {
4942+         all_ok = stbi__err("invalid filter","Corrupt PNG");
4943+         break;
4944+      }
4945+
4946+      /* if first row, use special filter that doesn't sample previous row */
4947+      if (j == 0) filter = first_row_filter[filter];
4948+
4949+      /* perform actual filtering */
4950+      switch (filter) {
4951+      case STBI__F_none:
4952+         memcpy(cur, raw, nk);
4953+         break;
4954+      case STBI__F_sub:
4955+         memcpy(cur, raw, filter_bytes);
4956+         for (k = filter_bytes; k < nk; ++k)
4957+            cur[k] = STBI__BYTECAST(raw[k] + cur[k-filter_bytes]);
4958+         break;
4959+      case STBI__F_up:
4960+         for (k = 0; k < nk; ++k)
4961+            cur[k] = STBI__BYTECAST(raw[k] + prior[k]);
4962+         break;
4963+      case STBI__F_avg:
4964+         for (k = 0; k < filter_bytes; ++k)
4965+            cur[k] = STBI__BYTECAST(raw[k] + (prior[k]>>1));
4966+         for (k = filter_bytes; k < nk; ++k)
4967+            cur[k] = STBI__BYTECAST(raw[k] + ((prior[k] + cur[k-filter_bytes])>>1));
4968+         break;
4969+      case STBI__F_paeth:
4970+         for (k = 0; k < filter_bytes; ++k)
4971+            cur[k] = STBI__BYTECAST(raw[k] + prior[k]); /* prior[k] == stbi__paeth(0,prior[k],0) */
4972+         for (k = filter_bytes; k < nk; ++k)
4973+            cur[k] = STBI__BYTECAST(raw[k] + stbi__paeth(cur[k-filter_bytes], prior[k], prior[k-filter_bytes]));
4974+         break;
4975+      case STBI__F_avg_first:
4976+         memcpy(cur, raw, filter_bytes);
4977+         for (k = filter_bytes; k < nk; ++k)
4978+            cur[k] = STBI__BYTECAST(raw[k] + (cur[k-filter_bytes] >> 1));
4979+         break;
4980+      }
4981+
4982+      raw += nk;
4983+
4984+      /* expand decoded bits in cur to dest, also adding an extra alpha channel if desired */
4985+      if (depth < 8) {
4986+         stbi_uc scale = (color == 0) ? stbi__depth_scale_table[depth] : 1; /* scale grayscale values to 0..255 range */
4987+         stbi_uc *in = cur;
4988+         stbi_uc *out = dest;
4989+         stbi_uc inb = 0;
4990+         stbi__uint32 nsmp = x*img_n;
4991+
4992+         /* expand bits to bytes first */
4993+         if (depth == 4) {
4994+            for (i=0; i < nsmp; ++i) {
4995+               if ((i & 1) == 0) inb = *in++;
4996+               *out++ = scale * (inb >> 4);
4997+               inb <<= 4;
4998+            }
4999+         } else if (depth == 2) {
5000+            for (i=0; i < nsmp; ++i) {
5001+               if ((i & 3) == 0) inb = *in++;
5002+               *out++ = scale * (inb >> 6);
5003+               inb <<= 2;
5004+            }
5005+         } else {
5006+            STBI_ASSERT(depth == 1);
5007+            for (i=0; i < nsmp; ++i) {
5008+               if ((i & 7) == 0) inb = *in++;
5009+               *out++ = scale * (inb >> 7);
5010+               inb <<= 1;
5011+            }
5012+         }
5013+
5014+         /* insert alpha=255 values if desired */
5015+         if (img_n != out_n)
5016+            stbi__create_png_alpha_expand8(dest, dest, x, img_n);
5017+      } else if (depth == 8) {
5018+         if (img_n == out_n)
5019+            memcpy(dest, cur, x*img_n);
5020+         else
5021+            stbi__create_png_alpha_expand8(dest, cur, x, img_n);
5022+      } else if (depth == 16) {
5023+         /* convert the image data from big-endian to platform-native */
5024+         stbi__uint16 *dest16 = (stbi__uint16*)dest;
5025+         stbi__uint32 nsmp = x*img_n;
5026+
5027+         if (img_n == out_n) {
5028+            for (i = 0; i < nsmp; ++i, ++dest16, cur += 2)
5029+               *dest16 = (cur[0] << 8) | cur[1];
5030+         } else {
5031+            STBI_ASSERT(img_n+1 == out_n);
5032+            if (img_n == 1) {
5033+               for (i = 0; i < x; ++i, dest16 += 2, cur += 2) {
5034+                  dest16[0] = (cur[0] << 8) | cur[1];
5035+                  dest16[1] = 0xffff;
5036+               }
5037+            } else {
5038+               STBI_ASSERT(img_n == 3);
5039+               for (i = 0; i < x; ++i, dest16 += 4, cur += 6) {
5040+                  dest16[0] = (cur[0] << 8) | cur[1];
5041+                  dest16[1] = (cur[2] << 8) | cur[3];
5042+                  dest16[2] = (cur[4] << 8) | cur[5];
5043+                  dest16[3] = 0xffff;
5044+               }
5045+            }
5046+         }
5047+      }
5048+   }
5049+
5050+   STBI_FREE(filter_buf);
5051+   if (!all_ok) return 0;
5052+
5053+   return 1;
5054+}
5055+
5056+static int stbi__create_png_image(stbi__png *a, stbi_uc *image_data, stbi__uint32 image_data_len, int out_n, int depth, int color, int interlaced)
5057+{
5058+   int bytes = (depth == 16 ? 2 : 1);
5059+   int out_bytes = out_n * bytes;
5060+   stbi_uc *final;
5061+   int p;
5062+   if (!interlaced)
5063+      return stbi__create_png_image_raw(a, image_data, image_data_len, out_n, a->s->img_x, a->s->img_y, depth, color);
5064+
5065+   /* de-interlacing */
5066+   final = (stbi_uc *) stbi__malloc_mad3(a->s->img_x, a->s->img_y, out_bytes, 0);
5067+   if (!final) return stbi__err("outofmem", "Out of memory");
5068+   for (p=0; p < 7; ++p) {
5069+      int xorig[] = { 0,4,0,2,0,1,0 };
5070+      int yorig[] = { 0,0,4,0,2,0,1 };
5071+      int xspc[]  = { 8,8,4,4,2,2,1 };
5072+      int yspc[]  = { 8,8,8,4,4,2,2 };
5073+      int i,j,x,y;
5074+      /* pass1_x[4] = 0, pass1_x[5] = 1, pass1_x[12] = 1 */
5075+      x = (a->s->img_x - xorig[p] + xspc[p]-1) / xspc[p];
5076+      y = (a->s->img_y - yorig[p] + yspc[p]-1) / yspc[p];
5077+      if (x && y) {
5078+         stbi__uint32 img_len = ((((a->s->img_n * x * depth) + 7) >> 3) + 1) * y;
5079+         if (!stbi__create_png_image_raw(a, image_data, image_data_len, out_n, x, y, depth, color)) {
5080+            STBI_FREE(final);
5081+            return 0;
5082+         }
5083+         for (j=0; j < y; ++j) {
5084+            for (i=0; i < x; ++i) {
5085+               int out_y = j*yspc[p]+yorig[p];
5086+               int out_x = i*xspc[p]+xorig[p];
5087+               memcpy(final + out_y*a->s->img_x*out_bytes + out_x*out_bytes,
5088+                      a->out + (j*x+i)*out_bytes, out_bytes);
5089+            }
5090+         }
5091+         STBI_FREE(a->out);
5092+         image_data += img_len;
5093+         image_data_len -= img_len;
5094+      }
5095+   }
5096+   a->out = final;
5097+
5098+   return 1;
5099+}
5100+
5101+static int stbi__compute_transparency(stbi__png *z, stbi_uc tc[3], int out_n)
5102+{
5103+   stbi__context *s = z->s;
5104+   stbi__uint32 i, pixel_count = s->img_x * s->img_y;
5105+   stbi_uc *p = z->out;
5106+
5107+   /*
5108+    * compute color-based transparency, assuming we've
5109+    * already got 255 as the alpha value in the output
5110+    */
5111+   STBI_ASSERT(out_n == 2 || out_n == 4);
5112+
5113+   if (out_n == 2) {
5114+      for (i=0; i < pixel_count; ++i) {
5115+         p[1] = (p[0] == tc[0] ? 0 : 255);
5116+         p += 2;
5117+      }
5118+   } else {
5119+      for (i=0; i < pixel_count; ++i) {
5120+         if (p[0] == tc[0] && p[1] == tc[1] && p[2] == tc[2])
5121+            p[3] = 0;
5122+         p += 4;
5123+      }
5124+   }
5125+   return 1;
5126+}
5127+
5128+static int stbi__compute_transparency16(stbi__png *z, stbi__uint16 tc[3], int out_n)
5129+{
5130+   stbi__context *s = z->s;
5131+   stbi__uint32 i, pixel_count = s->img_x * s->img_y;
5132+   stbi__uint16 *p = (stbi__uint16*) z->out;
5133+
5134+   /*
5135+    * compute color-based transparency, assuming we've
5136+    * already got 65535 as the alpha value in the output
5137+    */
5138+   STBI_ASSERT(out_n == 2 || out_n == 4);
5139+
5140+   if (out_n == 2) {
5141+      for (i = 0; i < pixel_count; ++i) {
5142+         p[1] = (p[0] == tc[0] ? 0 : 65535);
5143+         p += 2;
5144+      }
5145+   } else {
5146+      for (i = 0; i < pixel_count; ++i) {
5147+         if (p[0] == tc[0] && p[1] == tc[1] && p[2] == tc[2])
5148+            p[3] = 0;
5149+         p += 4;
5150+      }
5151+   }
5152+   return 1;
5153+}
5154+
5155+static int stbi__expand_png_palette(stbi__png *a, stbi_uc *palette, int len, int pal_img_n)
5156+{
5157+   stbi__uint32 i, pixel_count = a->s->img_x * a->s->img_y;
5158+   stbi_uc *p, *temp_out, *orig = a->out;
5159+
5160+   p = (stbi_uc *) stbi__malloc_mad2(pixel_count, pal_img_n, 0);
5161+   if (p == NULL) return stbi__err("outofmem", "Out of memory");
5162+
5163+   /* between here and free(out) below, exitting would leak */
5164+   temp_out = p;
5165+
5166+   if (pal_img_n == 3) {
5167+      for (i=0; i < pixel_count; ++i) {
5168+         int n = orig[i]*4;
5169+         p[0] = palette[n  ];
5170+         p[1] = palette[n+1];
5171+         p[2] = palette[n+2];
5172+         p += 3;
5173+      }
5174+   } else {
5175+      for (i=0; i < pixel_count; ++i) {
5176+         int n = orig[i]*4;
5177+         p[0] = palette[n  ];
5178+         p[1] = palette[n+1];
5179+         p[2] = palette[n+2];
5180+         p[3] = palette[n+3];
5181+         p += 4;
5182+      }
5183+   }
5184+   STBI_FREE(a->out);
5185+   a->out = temp_out;
5186+
5187+   STBI_NOTUSED(len);
5188+
5189+   return 1;
5190+}
5191+
5192+static int stbi__unpremultiply_on_load_global = 0;
5193+static int stbi__de_iphone_flag_global = 0;
5194+
5195+STBIDEF void stbi_set_unpremultiply_on_load(int flag_true_if_should_unpremultiply)
5196+{
5197+   stbi__unpremultiply_on_load_global = flag_true_if_should_unpremultiply;
5198+}
5199+
5200+STBIDEF void stbi_convert_iphone_png_to_rgb(int flag_true_if_should_convert)
5201+{
5202+   stbi__de_iphone_flag_global = flag_true_if_should_convert;
5203+}
5204+
5205+#ifndef STBI_THREAD_LOCAL
5206+#define stbi__unpremultiply_on_load  stbi__unpremultiply_on_load_global
5207+#define stbi__de_iphone_flag  stbi__de_iphone_flag_global
5208+#else
5209+static STBI_THREAD_LOCAL int stbi__unpremultiply_on_load_local, stbi__unpremultiply_on_load_set;
5210+static STBI_THREAD_LOCAL int stbi__de_iphone_flag_local, stbi__de_iphone_flag_set;
5211+
5212+STBIDEF void stbi_set_unpremultiply_on_load_thread(int flag_true_if_should_unpremultiply)
5213+{
5214+   stbi__unpremultiply_on_load_local = flag_true_if_should_unpremultiply;
5215+   stbi__unpremultiply_on_load_set = 1;
5216+}
5217+
5218+STBIDEF void stbi_convert_iphone_png_to_rgb_thread(int flag_true_if_should_convert)
5219+{
5220+   stbi__de_iphone_flag_local = flag_true_if_should_convert;
5221+   stbi__de_iphone_flag_set = 1;
5222+}
5223+
5224+#define stbi__unpremultiply_on_load  (stbi__unpremultiply_on_load_set           \
5225+                                       ? stbi__unpremultiply_on_load_local      \
5226+                                       : stbi__unpremultiply_on_load_global)
5227+#define stbi__de_iphone_flag  (stbi__de_iphone_flag_set                         \
5228+                                ? stbi__de_iphone_flag_local                    \
5229+                                : stbi__de_iphone_flag_global)
5230+#endif /* STBI_THREAD_LOCAL */
5231+
5232+static void stbi__de_iphone(stbi__png *z)
5233+{
5234+   stbi__context *s = z->s;
5235+   stbi__uint32 i, pixel_count = s->img_x * s->img_y;
5236+   stbi_uc *p = z->out;
5237+
5238+   if (s->img_out_n == 3) { /* convert bgr to rgb */
5239+      for (i=0; i < pixel_count; ++i) {
5240+         stbi_uc t = p[0];
5241+         p[0] = p[2];
5242+         p[2] = t;
5243+         p += 3;
5244+      }
5245+   } else {
5246+      STBI_ASSERT(s->img_out_n == 4);
5247+      if (stbi__unpremultiply_on_load) {
5248+         /* convert bgr to rgb and unpremultiply */
5249+         for (i=0; i < pixel_count; ++i) {
5250+            stbi_uc a = p[3];
5251+            stbi_uc t = p[0];
5252+            if (a) {
5253+               stbi_uc half = a / 2;
5254+               p[0] = (p[2] * 255 + half) / a;
5255+               p[1] = (p[1] * 255 + half) / a;
5256+               p[2] = ( t   * 255 + half) / a;
5257+            } else {
5258+               p[0] = p[2];
5259+               p[2] = t;
5260+            }
5261+            p += 4;
5262+         }
5263+      } else {
5264+         /* convert bgr to rgb */
5265+         for (i=0; i < pixel_count; ++i) {
5266+            stbi_uc t = p[0];
5267+            p[0] = p[2];
5268+            p[2] = t;
5269+            p += 4;
5270+         }
5271+      }
5272+   }
5273+}
5274+
5275+#define STBI__PNG_TYPE(a,b,c,d)  (((unsigned) (a) << 24) + ((unsigned) (b) << 16) + ((unsigned) (c) << 8) + (unsigned) (d))
5276+
5277+static int stbi__parse_png_file(stbi__png *z, int scan, int req_comp)
5278+{
5279+   stbi_uc palette[1024], pal_img_n=0;
5280+   stbi_uc has_trans=0, tc[3]={0};
5281+   stbi__uint16 tc16[3];
5282+   stbi__uint32 ioff=0, idata_limit=0, i, pal_len=0;
5283+   int first=1,k,interlace=0, color=0, is_iphone=0;
5284+   stbi__context *s = z->s;
5285+
5286+   z->expanded = NULL;
5287+   z->idata = NULL;
5288+   z->out = NULL;
5289+
5290+   if (!stbi__check_png_header(s)) return 0;
5291+
5292+   if (scan == STBI__SCAN_type) return 1;
5293+
5294+   for (;;) {
5295+      stbi__pngchunk c = stbi__get_chunk_header(s);
5296+      switch (c.type) {
5297+         case STBI__PNG_TYPE('C','g','B','I'):
5298+            is_iphone = 1;
5299+            stbi__skip(s, c.length);
5300+            break;
5301+         case STBI__PNG_TYPE('I','H','D','R'): {
5302+            int comp,filter;
5303+            if (!first) return stbi__err("multiple IHDR","Corrupt PNG");
5304+            first = 0;
5305+            if (c.length != 13) return stbi__err("bad IHDR len","Corrupt PNG");
5306+            s->img_x = stbi__get32be(s);
5307+            s->img_y = stbi__get32be(s);
5308+            if (s->img_y > STBI_MAX_DIMENSIONS) return stbi__err("too large","Very large image (corrupt?)");
5309+            if (s->img_x > STBI_MAX_DIMENSIONS) return stbi__err("too large","Very large image (corrupt?)");
5310+            z->depth = stbi__get8(s);  if (z->depth != 1 && z->depth != 2 && z->depth != 4 && z->depth != 8 && z->depth != 16)  return stbi__err("1/2/4/8/16-bit only","PNG not supported: 1/2/4/8/16-bit only");
5311+            color = stbi__get8(s);  if (color > 6)         return stbi__err("bad ctype","Corrupt PNG");
5312+            if (color == 3 && z->depth == 16)                  return stbi__err("bad ctype","Corrupt PNG");
5313+            if (color == 3) pal_img_n = 3; else if (color & 1) return stbi__err("bad ctype","Corrupt PNG");
5314+            comp  = stbi__get8(s);  if (comp) return stbi__err("bad comp method","Corrupt PNG");
5315+            filter= stbi__get8(s);  if (filter) return stbi__err("bad filter method","Corrupt PNG");
5316+            interlace = stbi__get8(s); if (interlace>1) return stbi__err("bad interlace method","Corrupt PNG");
5317+            if (!s->img_x || !s->img_y) return stbi__err("0-pixel image","Corrupt PNG");
5318+            if (!pal_img_n) {
5319+               s->img_n = (color & 2 ? 3 : 1) + (color & 4 ? 1 : 0);
5320+               if ((1 << 30) / s->img_x / s->img_n < s->img_y) return stbi__err("too large", "Image too large to decode");
5321+            } else {
5322+               /*
5323+                * if paletted, then pal_n is our final components, and
5324+                * img_n is # components to decompress/filter.
5325+                */
5326+               s->img_n = 1;
5327+               if ((1 << 30) / s->img_x / 4 < s->img_y) return stbi__err("too large","Corrupt PNG");
5328+            }
5329+            /* even with SCAN_header, have to scan to see if we have a tRNS */
5330+            break;
5331+         }
5332+
5333+         case STBI__PNG_TYPE('P','L','T','E'):  {
5334+            if (first) return stbi__err("first not IHDR", "Corrupt PNG");
5335+            if (c.length > 256*3) return stbi__err("invalid PLTE","Corrupt PNG");
5336+            pal_len = c.length / 3;
5337+            if (pal_len * 3 != c.length) return stbi__err("invalid PLTE","Corrupt PNG");
5338+            for (i=0; i < pal_len; ++i) {
5339+               palette[i*4+0] = stbi__get8(s);
5340+               palette[i*4+1] = stbi__get8(s);
5341+               palette[i*4+2] = stbi__get8(s);
5342+               palette[i*4+3] = 255;
5343+            }
5344+            break;
5345+         }
5346+
5347+         case STBI__PNG_TYPE('t','R','N','S'): {
5348+            if (first) return stbi__err("first not IHDR", "Corrupt PNG");
5349+            if (z->idata) return stbi__err("tRNS after IDAT","Corrupt PNG");
5350+            if (pal_img_n) {
5351+               if (scan == STBI__SCAN_header) { s->img_n = 4; return 1; }
5352+               if (pal_len == 0) return stbi__err("tRNS before PLTE","Corrupt PNG");
5353+               if (c.length > pal_len) return stbi__err("bad tRNS len","Corrupt PNG");
5354+               pal_img_n = 4;
5355+               for (i=0; i < c.length; ++i)
5356+                  palette[i*4+3] = stbi__get8(s);
5357+            } else {
5358+               if (!(s->img_n & 1)) return stbi__err("tRNS with alpha","Corrupt PNG");
5359+               if (c.length != (stbi__uint32) s->img_n*2) return stbi__err("bad tRNS len","Corrupt PNG");
5360+               has_trans = 1;
5361+               /* non-paletted with tRNS = constant alpha. if header-scanning, we can stop now. */
5362+               if (scan == STBI__SCAN_header) { ++s->img_n; return 1; }
5363+               if (z->depth == 16) {
5364+                  for (k = 0; k < s->img_n && k < 3; ++k) /* extra loop test to suppress false GCC warning */
5365+                     tc16[k] = (stbi__uint16)stbi__get16be(s); /* copy the values as-is */
5366+               } else {
5367+                  for (k = 0; k < s->img_n && k < 3; ++k)
5368+                     tc[k] = (stbi_uc)(stbi__get16be(s) & 255) * stbi__depth_scale_table[z->depth]; /* non 8-bit images will be larger */
5369+               }
5370+            }
5371+            break;
5372+         }
5373+
5374+         case STBI__PNG_TYPE('I','D','A','T'): {
5375+            if (first) return stbi__err("first not IHDR", "Corrupt PNG");
5376+            if (pal_img_n && !pal_len) return stbi__err("no PLTE","Corrupt PNG");
5377+            if (scan == STBI__SCAN_header) {
5378+               /* header scan definitely stops at first IDAT */
5379+               if (pal_img_n)
5380+                  s->img_n = pal_img_n;
5381+               return 1;
5382+            }
5383+            if (c.length > (1u << 30)) return stbi__err("IDAT size limit", "IDAT section larger than 2^30 bytes");
5384+            if ((int)(ioff + c.length) < (int)ioff) return 0;
5385+            if (ioff + c.length > idata_limit) {
5386+               stbi__uint32 idata_limit_old = idata_limit;
5387+               stbi_uc *p;
5388+               if (idata_limit == 0) idata_limit = c.length > 4096 ? c.length : 4096;
5389+               while (ioff + c.length > idata_limit)
5390+                  idata_limit *= 2;
5391+               STBI_NOTUSED(idata_limit_old);
5392+               p = (stbi_uc *) STBI_REALLOC_SIZED(z->idata, idata_limit_old, idata_limit); if (p == NULL) return stbi__err("outofmem", "Out of memory");
5393+               z->idata = p;
5394+            }
5395+            if (!stbi__getn(s, z->idata+ioff,c.length)) return stbi__err("outofdata","Corrupt PNG");
5396+            ioff += c.length;
5397+            break;
5398+         }
5399+
5400+         case STBI__PNG_TYPE('I','E','N','D'): {
5401+            stbi__uint32 raw_len, bpl;
5402+            if (first) return stbi__err("first not IHDR", "Corrupt PNG");
5403+            if (scan != STBI__SCAN_load) return 1;
5404+            if (z->idata == NULL) return stbi__err("no IDAT","Corrupt PNG");
5405+            /* initial guess for decoded data size to avoid unnecessary reallocs */
5406+            bpl = (s->img_x * z->depth + 7) / 8; /* bytes per line, per component */
5407+            raw_len = bpl * s->img_y * s->img_n /* pixels */ + s->img_y /* filter mode per row */;
5408+            z->expanded = (stbi_uc *) stbi_zlib_decode_malloc_guesssize_headerflag((char *) z->idata, ioff, raw_len, (int *) &raw_len, !is_iphone);
5409+            if (z->expanded == NULL) return 0; /* zlib should set error */
5410+            STBI_FREE(z->idata); z->idata = NULL;
5411+            if ((req_comp == s->img_n+1 && req_comp != 3 && !pal_img_n) || has_trans)
5412+               s->img_out_n = s->img_n+1;
5413+            else
5414+               s->img_out_n = s->img_n;
5415+            if (!stbi__create_png_image(z, z->expanded, raw_len, s->img_out_n, z->depth, color, interlace)) return 0;
5416+            if (has_trans) {
5417+               if (z->depth == 16) {
5418+                  if (!stbi__compute_transparency16(z, tc16, s->img_out_n)) return 0;
5419+               } else {
5420+                  if (!stbi__compute_transparency(z, tc, s->img_out_n)) return 0;
5421+               }
5422+            }
5423+            if (is_iphone && stbi__de_iphone_flag && s->img_out_n > 2)
5424+               stbi__de_iphone(z);
5425+            if (pal_img_n) {
5426+               /* pal_img_n == 3 or 4 */
5427+               s->img_n = pal_img_n; /* record the actual colors we had */
5428+               s->img_out_n = pal_img_n;
5429+               if (req_comp >= 3) s->img_out_n = req_comp;
5430+               if (!stbi__expand_png_palette(z, palette, pal_len, s->img_out_n))
5431+                  return 0;
5432+            } else if (has_trans) {
5433+               /* non-paletted image with tRNS -> source image has (constant) alpha */
5434+               ++s->img_n;
5435+            }
5436+            STBI_FREE(z->expanded); z->expanded = NULL;
5437+            /* end of PNG chunk, read and skip CRC */
5438+            stbi__get32be(s);
5439+            return 1;
5440+         }
5441+
5442+         default:
5443+            /* if critical, fail */
5444+            if (first) return stbi__err("first not IHDR", "Corrupt PNG");
5445+            if ((c.type & (1 << 29)) == 0) {
5446+               #ifndef STBI_NO_FAILURE_STRINGS
5447+               /* not threadsafe */
5448+               static char invalid_chunk[] = "XXXX PNG chunk not known";
5449+               invalid_chunk[0] = STBI__BYTECAST(c.type >> 24);
5450+               invalid_chunk[1] = STBI__BYTECAST(c.type >> 16);
5451+               invalid_chunk[2] = STBI__BYTECAST(c.type >>  8);
5452+               invalid_chunk[3] = STBI__BYTECAST(c.type >>  0);
5453+               #endif
5454+               return stbi__err(invalid_chunk, "PNG not supported: unknown PNG chunk type");
5455+            }
5456+            stbi__skip(s, c.length);
5457+            break;
5458+      }
5459+      /* end of PNG chunk, read and skip CRC */
5460+      stbi__get32be(s);
5461+   }
5462+}
5463+
5464+static void *stbi__do_png(stbi__png *p, int *x, int *y, int *n, int req_comp, stbi__result_info *ri)
5465+{
5466+   void *result=NULL;
5467+   if (req_comp < 0 || req_comp > 4) return stbi__errpuc("bad req_comp", "Internal error");
5468+   if (stbi__parse_png_file(p, STBI__SCAN_load, req_comp)) {
5469+      if (p->depth <= 8)
5470+         ri->bits_per_channel = 8;
5471+      else if (p->depth == 16)
5472+         ri->bits_per_channel = 16;
5473+      else
5474+         return stbi__errpuc("bad bits_per_channel", "PNG not supported: unsupported color depth");
5475+      result = p->out;
5476+      p->out = NULL;
5477+      if (req_comp && req_comp != p->s->img_out_n) {
5478+         if (ri->bits_per_channel == 8)
5479+            result = stbi__convert_format((unsigned char *) result, p->s->img_out_n, req_comp, p->s->img_x, p->s->img_y);
5480+         else
5481+            result = stbi__convert_format16((stbi__uint16 *) result, p->s->img_out_n, req_comp, p->s->img_x, p->s->img_y);
5482+         p->s->img_out_n = req_comp;
5483+         if (result == NULL) return result;
5484+      }
5485+      *x = p->s->img_x;
5486+      *y = p->s->img_y;
5487+      if (n) *n = p->s->img_n;
5488+   }
5489+   STBI_FREE(p->out);      p->out      = NULL;
5490+   STBI_FREE(p->expanded); p->expanded = NULL;
5491+   STBI_FREE(p->idata);    p->idata    = NULL;
5492+
5493+   return result;
5494+}
5495+
5496+static void *stbi__png_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri)
5497+{
5498+   stbi__png p;
5499+   p.s = s;
5500+   return stbi__do_png(&p, x,y,comp,req_comp, ri);
5501+}
5502+
5503+static int stbi__png_test(stbi__context *s)
5504+{
5505+   int r;
5506+   r = stbi__check_png_header(s);
5507+   stbi__rewind(s);
5508+   return r;
5509+}
5510+
5511+static int stbi__png_info_raw(stbi__png *p, int *x, int *y, int *comp)
5512+{
5513+   if (!stbi__parse_png_file(p, STBI__SCAN_header, 0)) {
5514+      stbi__rewind( p->s );
5515+      return 0;
5516+   }
5517+   if (x) *x = p->s->img_x;
5518+   if (y) *y = p->s->img_y;
5519+   if (comp) *comp = p->s->img_n;
5520+   return 1;
5521+}
5522+
5523+static int stbi__png_info(stbi__context *s, int *x, int *y, int *comp)
5524+{
5525+   stbi__png p;
5526+   p.s = s;
5527+   return stbi__png_info_raw(&p, x, y, comp);
5528+}
5529+
5530+static int stbi__png_is16(stbi__context *s)
5531+{
5532+   stbi__png p;
5533+   p.s = s;
5534+   if (!stbi__png_info_raw(&p, NULL, NULL, NULL))
5535+	   return 0;
5536+   if (p.depth != 16) {
5537+      stbi__rewind(p.s);
5538+      return 0;
5539+   }
5540+   return 1;
5541+}
5542+#endif
5543+
5544+/* Microsoft/Windows BMP image */
5545+
5546+#ifndef STBI_NO_BMP
5547+static int stbi__bmp_test_raw(stbi__context *s)
5548+{
5549+   int r;
5550+   int sz;
5551+   if (stbi__get8(s) != 'B') return 0;
5552+   if (stbi__get8(s) != 'M') return 0;
5553+   stbi__get32le(s); /* discard filesize */
5554+   stbi__get16le(s); /* discard reserved */
5555+   stbi__get16le(s); /* discard reserved */
5556+   stbi__get32le(s); /* discard data offset */
5557+   sz = stbi__get32le(s);
5558+   r = (sz == 12 || sz == 40 || sz == 56 || sz == 108 || sz == 124);
5559+   return r;
5560+}
5561+
5562+static int stbi__bmp_test(stbi__context *s)
5563+{
5564+   int r = stbi__bmp_test_raw(s);
5565+   stbi__rewind(s);
5566+   return r;
5567+}
5568+
5569+
5570+/* returns 0..31 for the highest set bit */
5571+static int stbi__high_bit(unsigned int z)
5572+{
5573+   int n=0;
5574+   if (z == 0) return -1;
5575+   if (z >= 0x10000) { n += 16; z >>= 16; }
5576+   if (z >= 0x00100) { n +=  8; z >>=  8; }
5577+   if (z >= 0x00010) { n +=  4; z >>=  4; }
5578+   if (z >= 0x00004) { n +=  2; z >>=  2; }
5579+   if (z >= 0x00002) { n +=  1;/* >>=  1;*/ }
5580+   return n;
5581+}
5582+
5583+static int stbi__bitcount(unsigned int a)
5584+{
5585+   a = (a & 0x55555555) + ((a >>  1) & 0x55555555); /* max 2 */
5586+   a = (a & 0x33333333) + ((a >>  2) & 0x33333333); /* max 4 */
5587+   a = (a + (a >> 4)) & 0x0f0f0f0f; /* max 8 per 4, now 8 bits */
5588+   a = (a + (a >> 8)); /* max 16 per 8 bits */
5589+   a = (a + (a >> 16)); /* max 32 per 8 bits */
5590+   return a & 0xff;
5591+}
5592+
5593+/*
5594+ * extract an arbitrarily-aligned N-bit value (N=bits)
5595+ * from v, and then make it 8-bits long and fractionally
5596+ * extend it to full full range.
5597+ */
5598+static int stbi__shiftsigned(unsigned int v, int shift, int bits)
5599+{
5600+   static unsigned int mul_table[9] = {
5601+      0,
5602+      0xff/*0b11111111*/, 0x55/*0b01010101*/, 0x49/*0b01001001*/, 0x11/*0b00010001*/,
5603+      0x21/*0b00100001*/, 0x41/*0b01000001*/, 0x81/*0b10000001*/, 0x01/*0b00000001*/,
5604+   };
5605+   static unsigned int shift_table[9] = {
5606+      0, 0,0,1,0,2,4,6,0,
5607+   };
5608+   if (shift < 0)
5609+      v <<= -shift;
5610+   else
5611+      v >>= shift;
5612+   STBI_ASSERT(v < 256);
5613+   v >>= (8-bits);
5614+   STBI_ASSERT(bits >= 0 && bits <= 8);
5615+   return (int) ((unsigned) v * mul_table[bits]) >> shift_table[bits];
5616+}
5617+
5618+typedef struct
5619+{
5620+   int bpp, offset, hsz;
5621+   unsigned int mr,mg,mb,ma, all_a;
5622+   int extra_read;
5623+} stbi__bmp_data;
5624+
5625+static int stbi__bmp_set_mask_defaults(stbi__bmp_data *info, int compress)
5626+{
5627+   /* BI_BITFIELDS specifies masks explicitly, don't override */
5628+   if (compress == 3)
5629+      return 1;
5630+
5631+   if (compress == 0) {
5632+      if (info->bpp == 16) {
5633+         info->mr = 31u << 10;
5634+         info->mg = 31u <<  5;
5635+         info->mb = 31u <<  0;
5636+      } else if (info->bpp == 32) {
5637+         info->mr = 0xffu << 16;
5638+         info->mg = 0xffu <<  8;
5639+         info->mb = 0xffu <<  0;
5640+         info->ma = 0xffu << 24;
5641+         info->all_a = 0; /* if all_a is 0 at end, then we loaded alpha channel but it was all 0 */
5642+      } else {
5643+         /* otherwise, use defaults, which is all-0 */
5644+         info->mr = info->mg = info->mb = info->ma = 0;
5645+      }
5646+      return 1;
5647+   }
5648+   return 0; /* error */
5649+}
5650+
5651+static void *stbi__bmp_parse_header(stbi__context *s, stbi__bmp_data *info)
5652+{
5653+   int hsz;
5654+   if (stbi__get8(s) != 'B' || stbi__get8(s) != 'M') return stbi__errpuc("not BMP", "Corrupt BMP");
5655+   stbi__get32le(s); /* discard filesize */
5656+   stbi__get16le(s); /* discard reserved */
5657+   stbi__get16le(s); /* discard reserved */
5658+   info->offset = stbi__get32le(s);
5659+   info->hsz = hsz = stbi__get32le(s);
5660+   info->mr = info->mg = info->mb = info->ma = 0;
5661+   info->extra_read = 14;
5662+
5663+   if (info->offset < 0) return stbi__errpuc("bad BMP", "bad BMP");
5664+
5665+   if (hsz != 12 && hsz != 40 && hsz != 56 && hsz != 108 && hsz != 124) return stbi__errpuc("unknown BMP", "BMP type not supported: unknown");
5666+   if (hsz == 12) {
5667+      s->img_x = stbi__get16le(s);
5668+      s->img_y = stbi__get16le(s);
5669+   } else {
5670+      s->img_x = stbi__get32le(s);
5671+      s->img_y = stbi__get32le(s);
5672+   }
5673+   if (stbi__get16le(s) != 1) return stbi__errpuc("bad BMP", "bad BMP");
5674+   info->bpp = stbi__get16le(s);
5675+   if (hsz != 12) {
5676+      int compress = stbi__get32le(s);
5677+      if (compress == 1 || compress == 2) return stbi__errpuc("BMP RLE", "BMP type not supported: RLE");
5678+      if (compress >= 4) return stbi__errpuc("BMP JPEG/PNG", "BMP type not supported: unsupported compression"); /* this includes PNG/JPEG modes */
5679+      if (compress == 3 && info->bpp != 16 && info->bpp != 32) return stbi__errpuc("bad BMP", "bad BMP"); /* bitfields requires 16 or 32 bits/pixel */
5680+      stbi__get32le(s); /* discard sizeof */
5681+      stbi__get32le(s); /* discard hres */
5682+      stbi__get32le(s); /* discard vres */
5683+      stbi__get32le(s); /* discard colorsused */
5684+      stbi__get32le(s); /* discard max important */
5685+      if (hsz == 40 || hsz == 56) {
5686+         if (hsz == 56) {
5687+            stbi__get32le(s);
5688+            stbi__get32le(s);
5689+            stbi__get32le(s);
5690+            stbi__get32le(s);
5691+         }
5692+         if (info->bpp == 16 || info->bpp == 32) {
5693+            if (compress == 0) {
5694+               stbi__bmp_set_mask_defaults(info, compress);
5695+            } else if (compress == 3) {
5696+               info->mr = stbi__get32le(s);
5697+               info->mg = stbi__get32le(s);
5698+               info->mb = stbi__get32le(s);
5699+               info->extra_read += 12;
5700+               /* not documented, but generated by photoshop and handled by mspaint */
5701+               if (info->mr == info->mg && info->mg == info->mb) {
5702+                  /* ?!?!? */
5703+                  return stbi__errpuc("bad BMP", "bad BMP");
5704+               }
5705+            } else
5706+               return stbi__errpuc("bad BMP", "bad BMP");
5707+         }
5708+      } else {
5709+         /* V4/V5 header */
5710+         int i;
5711+         if (hsz != 108 && hsz != 124)
5712+            return stbi__errpuc("bad BMP", "bad BMP");
5713+         info->mr = stbi__get32le(s);
5714+         info->mg = stbi__get32le(s);
5715+         info->mb = stbi__get32le(s);
5716+         info->ma = stbi__get32le(s);
5717+         if (compress != 3) /* override mr/mg/mb unless in BI_BITFIELDS mode, as per docs */
5718+            stbi__bmp_set_mask_defaults(info, compress);
5719+         stbi__get32le(s); /* discard color space */
5720+         for (i=0; i < 12; ++i)
5721+            stbi__get32le(s); /* discard color space parameters */
5722+         if (hsz == 124) {
5723+            stbi__get32le(s); /* discard rendering intent */
5724+            stbi__get32le(s); /* discard offset of profile data */
5725+            stbi__get32le(s); /* discard size of profile data */
5726+            stbi__get32le(s); /* discard reserved */
5727+         }
5728+      }
5729+   }
5730+   return (void *) 1;
5731+}
5732+
5733+
5734+static void *stbi__bmp_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri)
5735+{
5736+   stbi_uc *out;
5737+   unsigned int mr=0,mg=0,mb=0,ma=0, all_a;
5738+   stbi_uc pal[256][4];
5739+   int psize=0,i,j,width;
5740+   int flip_vertically, pad, target;
5741+   stbi__bmp_data info;
5742+   STBI_NOTUSED(ri);
5743+
5744+   info.all_a = 255;
5745+   if (stbi__bmp_parse_header(s, &info) == NULL)
5746+      return NULL; /* error code already set */
5747+
5748+   flip_vertically = ((int) s->img_y) > 0;
5749+   s->img_y = abs((int) s->img_y);
5750+
5751+   if (s->img_y > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
5752+   if (s->img_x > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
5753+
5754+   mr = info.mr;
5755+   mg = info.mg;
5756+   mb = info.mb;
5757+   ma = info.ma;
5758+   all_a = info.all_a;
5759+
5760+   if (info.hsz == 12) {
5761+      if (info.bpp < 24)
5762+         psize = (info.offset - info.extra_read - 24) / 3;
5763+   } else {
5764+      if (info.bpp < 16)
5765+         psize = (info.offset - info.extra_read - info.hsz) >> 2;
5766+   }
5767+   if (psize == 0) {
5768+      /*
5769+       * accept some number of extra bytes after the header, but if the offset points either to before
5770+       * the header ends or implies a large amount of extra data, reject the file as malformed
5771+       */
5772+      int bytes_read_so_far = s->callback_already_read + (int)(s->img_buffer - s->img_buffer_original);
5773+      int header_limit = 1024; /* max we actually read is below 256 bytes currently. */
5774+      int extra_data_limit = 256*4; /* what ordinarily goes here is a palette; 256 entries*4 bytes is its max size. */
5775+      if (bytes_read_so_far <= 0 || bytes_read_so_far > header_limit) {
5776+         return stbi__errpuc("bad header", "Corrupt BMP");
5777+      }
5778+      /*
5779+       * we established that bytes_read_so_far is positive and sensible.
5780+       * the first half of this test rejects offsets that are either too small positives, or
5781+       * negative, and guarantees that info.offset >= bytes_read_so_far > 0. this in turn
5782+       * ensures the number computed in the second half of the test can't overflow.
5783+       */
5784+      if (info.offset < bytes_read_so_far || info.offset - bytes_read_so_far > extra_data_limit) {
5785+         return stbi__errpuc("bad offset", "Corrupt BMP");
5786+      } else {
5787+         stbi__skip(s, info.offset - bytes_read_so_far);
5788+      }
5789+   }
5790+
5791+   if (info.bpp == 24 && ma == 0xff000000)
5792+      s->img_n = 3;
5793+   else
5794+      s->img_n = ma ? 4 : 3;
5795+   if (req_comp && req_comp >= 3) /* we can directly decode 3 or 4 */
5796+      target = req_comp;
5797+   else
5798+      target = s->img_n; /* if they want monochrome, we'll post-convert */
5799+
5800+   /* sanity-check size */
5801+   if (!stbi__mad3sizes_valid(target, s->img_x, s->img_y, 0))
5802+      return stbi__errpuc("too large", "Corrupt BMP");
5803+
5804+   out = (stbi_uc *) stbi__malloc_mad3(target, s->img_x, s->img_y, 0);
5805+   if (!out) return stbi__errpuc("outofmem", "Out of memory");
5806+   if (info.bpp < 16) {
5807+      int z=0;
5808+      if (psize == 0 || psize > 256) { STBI_FREE(out); return stbi__errpuc("invalid", "Corrupt BMP"); }
5809+      for (i=0; i < psize; ++i) {
5810+         pal[i][2] = stbi__get8(s);
5811+         pal[i][1] = stbi__get8(s);
5812+         pal[i][0] = stbi__get8(s);
5813+         if (info.hsz != 12) stbi__get8(s);
5814+         pal[i][3] = 255;
5815+      }
5816+      stbi__skip(s, info.offset - info.extra_read - info.hsz - psize * (info.hsz == 12 ? 3 : 4));
5817+      if (info.bpp == 1) width = (s->img_x + 7) >> 3;
5818+      else if (info.bpp == 4) width = (s->img_x + 1) >> 1;
5819+      else if (info.bpp == 8) width = s->img_x;
5820+      else { STBI_FREE(out); return stbi__errpuc("bad bpp", "Corrupt BMP"); }
5821+      pad = (-width)&3;
5822+      if (info.bpp == 1) {
5823+         for (j=0; j < (int) s->img_y; ++j) {
5824+            int bit_offset = 7, v = stbi__get8(s);
5825+            for (i=0; i < (int) s->img_x; ++i) {
5826+               int color = (v>>bit_offset)&0x1;
5827+               out[z++] = pal[color][0];
5828+               out[z++] = pal[color][1];
5829+               out[z++] = pal[color][2];
5830+               if (target == 4) out[z++] = 255;
5831+               if (i+1 == (int) s->img_x) break;
5832+               if((--bit_offset) < 0) {
5833+                  bit_offset = 7;
5834+                  v = stbi__get8(s);
5835+               }
5836+            }
5837+            stbi__skip(s, pad);
5838+         }
5839+      } else {
5840+         for (j=0; j < (int) s->img_y; ++j) {
5841+            for (i=0; i < (int) s->img_x; i += 2) {
5842+               int v=stbi__get8(s),v2=0;
5843+               if (info.bpp == 4) {
5844+                  v2 = v & 15;
5845+                  v >>= 4;
5846+               }
5847+               out[z++] = pal[v][0];
5848+               out[z++] = pal[v][1];
5849+               out[z++] = pal[v][2];
5850+               if (target == 4) out[z++] = 255;
5851+               if (i+1 == (int) s->img_x) break;
5852+               v = (info.bpp == 8) ? stbi__get8(s) : v2;
5853+               out[z++] = pal[v][0];
5854+               out[z++] = pal[v][1];
5855+               out[z++] = pal[v][2];
5856+               if (target == 4) out[z++] = 255;
5857+            }
5858+            stbi__skip(s, pad);
5859+         }
5860+      }
5861+   } else {
5862+      int rshift=0,gshift=0,bshift=0,ashift=0,rcount=0,gcount=0,bcount=0,acount=0;
5863+      int z = 0;
5864+      int easy=0;
5865+      stbi__skip(s, info.offset - info.extra_read - info.hsz);
5866+      if (info.bpp == 24) width = 3 * s->img_x;
5867+      else if (info.bpp == 16) width = 2*s->img_x;
5868+      else /* bpp = 32 and pad = 0 */ width=0;
5869+      pad = (-width) & 3;
5870+      if (info.bpp == 24) {
5871+         easy = 1;
5872+      } else if (info.bpp == 32) {
5873+         if (mb == 0xff && mg == 0xff00 && mr == 0x00ff0000 && ma == 0xff000000)
5874+            easy = 2;
5875+      }
5876+      if (!easy) {
5877+         if (!mr || !mg || !mb) { STBI_FREE(out); return stbi__errpuc("bad masks", "Corrupt BMP"); }
5878+         /* right shift amt to put high bit in position #7 */
5879+         rshift = stbi__high_bit(mr)-7; rcount = stbi__bitcount(mr);
5880+         gshift = stbi__high_bit(mg)-7; gcount = stbi__bitcount(mg);
5881+         bshift = stbi__high_bit(mb)-7; bcount = stbi__bitcount(mb);
5882+         ashift = stbi__high_bit(ma)-7; acount = stbi__bitcount(ma);
5883+         if (rcount > 8 || gcount > 8 || bcount > 8 || acount > 8) { STBI_FREE(out); return stbi__errpuc("bad masks", "Corrupt BMP"); }
5884+      }
5885+      for (j=0; j < (int) s->img_y; ++j) {
5886+         if (easy) {
5887+            for (i=0; i < (int) s->img_x; ++i) {
5888+               unsigned char a;
5889+               out[z+2] = stbi__get8(s);
5890+               out[z+1] = stbi__get8(s);
5891+               out[z+0] = stbi__get8(s);
5892+               z += 3;
5893+               a = (easy == 2 ? stbi__get8(s) : 255);
5894+               all_a |= a;
5895+               if (target == 4) out[z++] = a;
5896+            }
5897+         } else {
5898+            int bpp = info.bpp;
5899+            for (i=0; i < (int) s->img_x; ++i) {
5900+               stbi__uint32 v = (bpp == 16 ? (stbi__uint32) stbi__get16le(s) : stbi__get32le(s));
5901+               unsigned int a;
5902+               out[z++] = STBI__BYTECAST(stbi__shiftsigned(v & mr, rshift, rcount));
5903+               out[z++] = STBI__BYTECAST(stbi__shiftsigned(v & mg, gshift, gcount));
5904+               out[z++] = STBI__BYTECAST(stbi__shiftsigned(v & mb, bshift, bcount));
5905+               a = (ma ? stbi__shiftsigned(v & ma, ashift, acount) : 255);
5906+               all_a |= a;
5907+               if (target == 4) out[z++] = STBI__BYTECAST(a);
5908+            }
5909+         }
5910+         stbi__skip(s, pad);
5911+      }
5912+   }
5913+
5914+   /* if alpha channel is all 0s, replace with all 255s */
5915+   if (target == 4 && all_a == 0)
5916+      for (i=4*s->img_x*s->img_y-1; i >= 0; i -= 4)
5917+         out[i] = 255;
5918+
5919+   if (flip_vertically) {
5920+      stbi_uc t;
5921+      for (j=0; j < (int) s->img_y>>1; ++j) {
5922+         stbi_uc *p1 = out +      j     *s->img_x*target;
5923+         stbi_uc *p2 = out + (s->img_y-1-j)*s->img_x*target;
5924+         for (i=0; i < (int) s->img_x*target; ++i) {
5925+            t = p1[i]; p1[i] = p2[i]; p2[i] = t;
5926+         }
5927+      }
5928+   }
5929+
5930+   if (req_comp && req_comp != target) {
5931+      out = stbi__convert_format(out, target, req_comp, s->img_x, s->img_y);
5932+      if (out == NULL) return out; /* stbi__convert_format frees input on failure */
5933+   }
5934+
5935+   *x = s->img_x;
5936+   *y = s->img_y;
5937+   if (comp) *comp = s->img_n;
5938+   return out;
5939+}
5940+#endif
5941+
5942+/*
5943+ * Targa Truevision - TGA
5944+ * by Jonathan Dummer
5945+ */
5946+#ifndef STBI_NO_TGA
5947+/* returns STBI_rgb or whatever, 0 on error */
5948+static int stbi__tga_get_comp(int bits_per_pixel, int is_grey, int* is_rgb16)
5949+{
5950+   /* only RGB or RGBA (incl. 16bit) or grey allowed */
5951+   if (is_rgb16) *is_rgb16 = 0;
5952+   switch(bits_per_pixel) {
5953+      case 8:  return STBI_grey;
5954+      case 16: if(is_grey) return STBI_grey_alpha;
5955+               /* fallthrough */
5956+      case 15: if(is_rgb16) *is_rgb16 = 1;
5957+               return STBI_rgb;
5958+      case 24: /* fallthrough */
5959+      case 32: return bits_per_pixel/8;
5960+      default: return 0;
5961+   }
5962+}
5963+
5964+static int stbi__tga_info(stbi__context *s, int *x, int *y, int *comp)
5965+{
5966+    int tga_w, tga_h, tga_comp, tga_image_type, tga_bits_per_pixel, tga_colormap_bpp;
5967+    int sz, tga_colormap_type;
5968+    stbi__get8(s); /* discard Offset */
5969+    tga_colormap_type = stbi__get8(s); /* colormap type */
5970+    if( tga_colormap_type > 1 ) {
5971+        stbi__rewind(s);
5972+        return 0; /* only RGB or indexed allowed */
5973+    }
5974+    tga_image_type = stbi__get8(s); /* image type */
5975+    if ( tga_colormap_type == 1 ) { /* colormapped (paletted) image */
5976+        if (tga_image_type != 1 && tga_image_type != 9) {
5977+            stbi__rewind(s);
5978+            return 0;
5979+        }
5980+        stbi__skip(s,4); /* skip index of first colormap entry and number of entries */
5981+        sz = stbi__get8(s); /*   check bits per palette color entry */
5982+        if ( (sz != 8) && (sz != 15) && (sz != 16) && (sz != 24) && (sz != 32) ) {
5983+            stbi__rewind(s);
5984+            return 0;
5985+        }
5986+        stbi__skip(s,4); /* skip image x and y origin */
5987+        tga_colormap_bpp = sz;
5988+    } else { /* "normal" image w/o colormap - only RGB or grey allowed, +/- RLE */
5989+        if ( (tga_image_type != 2) && (tga_image_type != 3) && (tga_image_type != 10) && (tga_image_type != 11) ) {
5990+            stbi__rewind(s);
5991+            return 0; /* only RGB or grey allowed, +/- RLE */
5992+        }
5993+        stbi__skip(s,9); /* skip colormap specification and image x/y origin */
5994+        tga_colormap_bpp = 0;
5995+    }
5996+    tga_w = stbi__get16le(s);
5997+    if( tga_w < 1 ) {
5998+        stbi__rewind(s);
5999+        return 0; /* test width */
6000+    }
6001+    tga_h = stbi__get16le(s);
6002+    if( tga_h < 1 ) {
6003+        stbi__rewind(s);
6004+        return 0; /* test height */
6005+    }
6006+    tga_bits_per_pixel = stbi__get8(s); /* bits per pixel */
6007+    stbi__get8(s); /* ignore alpha bits */
6008+    if (tga_colormap_bpp != 0) {
6009+        if((tga_bits_per_pixel != 8) && (tga_bits_per_pixel != 16)) {
6010+            /*
6011+             * when using a colormap, tga_bits_per_pixel is the size of the indexes
6012+             * I don't think anything but 8 or 16bit indexes makes sense
6013+             */
6014+            stbi__rewind(s);
6015+            return 0;
6016+        }
6017+        tga_comp = stbi__tga_get_comp(tga_colormap_bpp, 0, NULL);
6018+    } else {
6019+        tga_comp = stbi__tga_get_comp(tga_bits_per_pixel, (tga_image_type == 3) || (tga_image_type == 11), NULL);
6020+    }
6021+    if(!tga_comp) {
6022+      stbi__rewind(s);
6023+      return 0;
6024+    }
6025+    if (x) *x = tga_w;
6026+    if (y) *y = tga_h;
6027+    if (comp) *comp = tga_comp;
6028+    return 1; /* seems to have passed everything */
6029+}
6030+
6031+static int stbi__tga_test(stbi__context *s)
6032+{
6033+   int res = 0;
6034+   int sz, tga_color_type;
6035+   stbi__get8(s); /*   discard Offset */
6036+   tga_color_type = stbi__get8(s); /*   color type */
6037+   if ( tga_color_type > 1 ) goto errorEnd; /*   only RGB or indexed allowed */
6038+   sz = stbi__get8(s); /*   image type */
6039+   if ( tga_color_type == 1 ) { /* colormapped (paletted) image */
6040+      if (sz != 1 && sz != 9) goto errorEnd; /* colortype 1 demands image type 1 or 9 */
6041+      stbi__skip(s,4); /* skip index of first colormap entry and number of entries */
6042+      sz = stbi__get8(s); /*   check bits per palette color entry */
6043+      if ( (sz != 8) && (sz != 15) && (sz != 16) && (sz != 24) && (sz != 32) ) goto errorEnd;
6044+      stbi__skip(s,4); /* skip image x and y origin */
6045+   } else { /* "normal" image w/o colormap */
6046+      if ( (sz != 2) && (sz != 3) && (sz != 10) && (sz != 11) ) goto errorEnd; /* only RGB or grey allowed, +/- RLE */
6047+      stbi__skip(s,9); /* skip colormap specification and image x/y origin */
6048+   }
6049+   if ( stbi__get16le(s) < 1 ) goto errorEnd; /*   test width */
6050+   if ( stbi__get16le(s) < 1 ) goto errorEnd; /*   test height */
6051+   sz = stbi__get8(s); /*   bits per pixel */
6052+   if ( (tga_color_type == 1) && (sz != 8) && (sz != 16) ) goto errorEnd; /* for colormapped images, bpp is size of an index */
6053+   if ( (sz != 8) && (sz != 15) && (sz != 16) && (sz != 24) && (sz != 32) ) goto errorEnd;
6054+
6055+   res = 1; /* if we got this far, everything's good and we can return 1 instead of 0 */
6056+
6057+errorEnd:
6058+   stbi__rewind(s);
6059+   return res;
6060+}
6061+
6062+/* read 16bit value and convert to 24bit RGB */
6063+static void stbi__tga_read_rgb16(stbi__context *s, stbi_uc* out)
6064+{
6065+   stbi__uint16 px = (stbi__uint16)stbi__get16le(s);
6066+   stbi__uint16 fiveBitMask = 31;
6067+   /* we have 3 channels with 5bits each */
6068+   int r = (px >> 10) & fiveBitMask;
6069+   int g = (px >> 5) & fiveBitMask;
6070+   int b = px & fiveBitMask;
6071+   /* Note that this saves the data in RGB(A) order, so it doesn't need to be swapped later */
6072+   out[0] = (stbi_uc)((r * 255)/31);
6073+   out[1] = (stbi_uc)((g * 255)/31);
6074+   out[2] = (stbi_uc)((b * 255)/31);
6075+
6076+   /*
6077+    * some people claim that the most significant bit might be used for alpha
6078+    * (possibly if an alpha-bit is set in the "image descriptor byte")
6079+    * but that only made 16bit test images completely translucent..
6080+    * so let's treat all 15 and 16bit TGAs as RGB with no alpha.
6081+    */
6082+}
6083+
6084+static void *stbi__tga_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri)
6085+{
6086+   /*   read in the TGA header stuff */
6087+   int tga_offset = stbi__get8(s);
6088+   int tga_indexed = stbi__get8(s);
6089+   int tga_image_type = stbi__get8(s);
6090+   int tga_is_RLE = 0;
6091+   int tga_palette_start = stbi__get16le(s);
6092+   int tga_palette_len = stbi__get16le(s);
6093+   int tga_palette_bits = stbi__get8(s);
6094+   int tga_x_origin = stbi__get16le(s);
6095+   int tga_y_origin = stbi__get16le(s);
6096+   int tga_width = stbi__get16le(s);
6097+   int tga_height = stbi__get16le(s);
6098+   int tga_bits_per_pixel = stbi__get8(s);
6099+   int tga_comp, tga_rgb16=0;
6100+   int tga_inverted = stbi__get8(s);
6101+   /*
6102+    * int tga_alpha_bits = tga_inverted & 15; // the 4 lowest bits - unused (useless?)
6103+    *   image data
6104+    */
6105+   unsigned char *tga_data;
6106+   unsigned char *tga_palette = NULL;
6107+   int i, j;
6108+   unsigned char raw_data[4] = {0};
6109+   int RLE_count = 0;
6110+   int RLE_repeating = 0;
6111+   int read_next_pixel = 1;
6112+   STBI_NOTUSED(ri);
6113+   STBI_NOTUSED(tga_x_origin); /* @TODO */
6114+   STBI_NOTUSED(tga_y_origin); /* @TODO */
6115+
6116+   if (tga_height > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
6117+   if (tga_width > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
6118+
6119+   /*   do a tiny bit of precessing */
6120+   if ( tga_image_type >= 8 )
6121+   {
6122+      tga_image_type -= 8;
6123+      tga_is_RLE = 1;
6124+   }
6125+   tga_inverted = 1 - ((tga_inverted >> 5) & 1);
6126+
6127+   /*   If I'm paletted, then I'll use the number of bits from the palette */
6128+   if ( tga_indexed ) tga_comp = stbi__tga_get_comp(tga_palette_bits, 0, &tga_rgb16);
6129+   else tga_comp = stbi__tga_get_comp(tga_bits_per_pixel, (tga_image_type == 3), &tga_rgb16);
6130+
6131+   if(!tga_comp) /* shouldn't really happen, stbi__tga_test() should have ensured basic consistency */
6132+      return stbi__errpuc("bad format", "Can't find out TGA pixelformat");
6133+
6134+   /*   tga info */
6135+   *x = tga_width;
6136+   *y = tga_height;
6137+   if (comp) *comp = tga_comp;
6138+
6139+   if (!stbi__mad3sizes_valid(tga_width, tga_height, tga_comp, 0))
6140+      return stbi__errpuc("too large", "Corrupt TGA");
6141+
6142+   tga_data = (unsigned char*)stbi__malloc_mad3(tga_width, tga_height, tga_comp, 0);
6143+   if (!tga_data) return stbi__errpuc("outofmem", "Out of memory");
6144+
6145+   /* skip to the data's starting position (offset usually = 0) */
6146+   stbi__skip(s, tga_offset );
6147+
6148+   if ( !tga_indexed && !tga_is_RLE && !tga_rgb16 ) {
6149+      for (i=0; i < tga_height; ++i) {
6150+         int row = tga_inverted ? tga_height -i - 1 : i;
6151+         stbi_uc *tga_row = tga_data + row*tga_width*tga_comp;
6152+         stbi__getn(s, tga_row, tga_width * tga_comp);
6153+      }
6154+   } else  {
6155+      /*   do I need to load a palette? */
6156+      if ( tga_indexed)
6157+      {
6158+         if (tga_palette_len == 0) {  /* you have to have at least one entry! */
6159+            STBI_FREE(tga_data);
6160+            return stbi__errpuc("bad palette", "Corrupt TGA");
6161+         }
6162+
6163+         /*   any data to skip? (offset usually = 0) */
6164+         stbi__skip(s, tga_palette_start );
6165+         /*   load the palette */
6166+         tga_palette = (unsigned char*)stbi__malloc_mad2(tga_palette_len, tga_comp, 0);
6167+         if (!tga_palette) {
6168+            STBI_FREE(tga_data);
6169+            return stbi__errpuc("outofmem", "Out of memory");
6170+         }
6171+         if (tga_rgb16) {
6172+            stbi_uc *pal_entry = tga_palette;
6173+            STBI_ASSERT(tga_comp == STBI_rgb);
6174+            for (i=0; i < tga_palette_len; ++i) {
6175+               stbi__tga_read_rgb16(s, pal_entry);
6176+               pal_entry += tga_comp;
6177+            }
6178+         } else if (!stbi__getn(s, tga_palette, tga_palette_len * tga_comp)) {
6179+               STBI_FREE(tga_data);
6180+               STBI_FREE(tga_palette);
6181+               return stbi__errpuc("bad palette", "Corrupt TGA");
6182+         }
6183+      }
6184+      /*   load the data */
6185+      for (i=0; i < tga_width * tga_height; ++i)
6186+      {
6187+         /*   if I'm in RLE mode, do I need to get a RLE stbi__pngchunk? */
6188+         if ( tga_is_RLE )
6189+         {
6190+            if ( RLE_count == 0 )
6191+            {
6192+               /*   yep, get the next byte as a RLE command */
6193+               int RLE_cmd = stbi__get8(s);
6194+               RLE_count = 1 + (RLE_cmd & 127);
6195+               RLE_repeating = RLE_cmd >> 7;
6196+               read_next_pixel = 1;
6197+            } else if ( !RLE_repeating )
6198+            {
6199+               read_next_pixel = 1;
6200+            }
6201+         } else
6202+         {
6203+            read_next_pixel = 1;
6204+         }
6205+         /*   OK, if I need to read a pixel, do it now */
6206+         if ( read_next_pixel )
6207+         {
6208+            /*   load however much data we did have */
6209+            if ( tga_indexed )
6210+            {
6211+               /* read in index, then perform the lookup */
6212+               int pal_idx = (tga_bits_per_pixel == 8) ? stbi__get8(s) : stbi__get16le(s);
6213+               if ( pal_idx >= tga_palette_len ) {
6214+                  /* invalid index */
6215+                  pal_idx = 0;
6216+               }
6217+               pal_idx *= tga_comp;
6218+               for (j = 0; j < tga_comp; ++j) {
6219+                  raw_data[j] = tga_palette[pal_idx+j];
6220+               }
6221+            } else if(tga_rgb16) {
6222+               STBI_ASSERT(tga_comp == STBI_rgb);
6223+               stbi__tga_read_rgb16(s, raw_data);
6224+            } else {
6225+               /*   read in the data raw */
6226+               for (j = 0; j < tga_comp; ++j) {
6227+                  raw_data[j] = stbi__get8(s);
6228+               }
6229+            }
6230+            /*   clear the reading flag for the next pixel */
6231+            read_next_pixel = 0;
6232+         } /* end of reading a pixel */
6233+
6234+         /* copy data */
6235+         for (j = 0; j < tga_comp; ++j)
6236+           tga_data[i*tga_comp+j] = raw_data[j];
6237+
6238+         /*   in case we're in RLE mode, keep counting down */
6239+         --RLE_count;
6240+      }
6241+      /*   do I need to invert the image? */
6242+      if ( tga_inverted )
6243+      {
6244+         for (j = 0; j*2 < tga_height; ++j)
6245+         {
6246+            int index1 = j * tga_width * tga_comp;
6247+            int index2 = (tga_height - 1 - j) * tga_width * tga_comp;
6248+            for (i = tga_width * tga_comp; i > 0; --i)
6249+            {
6250+               unsigned char temp = tga_data[index1];
6251+               tga_data[index1] = tga_data[index2];
6252+               tga_data[index2] = temp;
6253+               ++index1;
6254+               ++index2;
6255+            }
6256+         }
6257+      }
6258+      /*   clear my palette, if I had one */
6259+      if ( tga_palette != NULL )
6260+      {
6261+         STBI_FREE( tga_palette );
6262+      }
6263+   }
6264+
6265+   /* swap RGB - if the source data was RGB16, it already is in the right order */
6266+   if (tga_comp >= 3 && !tga_rgb16)
6267+   {
6268+      unsigned char* tga_pixel = tga_data;
6269+      for (i=0; i < tga_width * tga_height; ++i)
6270+      {
6271+         unsigned char temp = tga_pixel[0];
6272+         tga_pixel[0] = tga_pixel[2];
6273+         tga_pixel[2] = temp;
6274+         tga_pixel += tga_comp;
6275+      }
6276+   }
6277+
6278+   /* convert to target component count */
6279+   if (req_comp && req_comp != tga_comp)
6280+      tga_data = stbi__convert_format(tga_data, tga_comp, req_comp, tga_width, tga_height);
6281+
6282+   /*
6283+    *   the things I do to get rid of an error message, and yet keep
6284+    *   Microsoft's C compilers happy... [8^(
6285+    */
6286+   tga_palette_start = tga_palette_len = tga_palette_bits =
6287+         tga_x_origin = tga_y_origin = 0;
6288+   STBI_NOTUSED(tga_palette_start);
6289+   /*   OK, done */
6290+   return tga_data;
6291+}
6292+#endif
6293+
6294+/***************************************************************************************************/
6295+/* Photoshop PSD loader -- PD by Thatcher Ulrich, integration by Nicolas Schulz, tweaked by STB */
6296+
6297+#ifndef STBI_NO_PSD
6298+static int stbi__psd_test(stbi__context *s)
6299+{
6300+   int r = (stbi__get32be(s) == 0x38425053);
6301+   stbi__rewind(s);
6302+   return r;
6303+}
6304+
6305+static int stbi__psd_decode_rle(stbi__context *s, stbi_uc *p, int pixelCount)
6306+{
6307+   int count, nleft, len;
6308+
6309+   count = 0;
6310+   while ((nleft = pixelCount - count) > 0) {
6311+      len = stbi__get8(s);
6312+      if (len == 128) {
6313+         /* No-op. */
6314+      } else if (len < 128) {
6315+         /* Copy next len+1 bytes literally. */
6316+         len++;
6317+         if (len > nleft) return 0; /* corrupt data */
6318+         count += len;
6319+         while (len) {
6320+            *p = stbi__get8(s);
6321+            p += 4;
6322+            len--;
6323+         }
6324+      } else if (len > 128) {
6325+         stbi_uc   val;
6326+         /*
6327+          * Next -len+1 bytes in the dest are replicated from next source byte.
6328+          * (Interpret len as a negative 8-bit int.)
6329+          */
6330+         len = 257 - len;
6331+         if (len > nleft) return 0; /* corrupt data */
6332+         val = stbi__get8(s);
6333+         count += len;
6334+         while (len) {
6335+            *p = val;
6336+            p += 4;
6337+            len--;
6338+         }
6339+      }
6340+   }
6341+
6342+   return 1;
6343+}
6344+
6345+static void *stbi__psd_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri, int bpc)
6346+{
6347+   int pixelCount;
6348+   int channelCount, compression;
6349+   int channel, i;
6350+   int bitdepth;
6351+   int w,h;
6352+   stbi_uc *out;
6353+   STBI_NOTUSED(ri);
6354+
6355+   /* Check identifier */
6356+   if (stbi__get32be(s) != 0x38425053) /* "8BPS" */
6357+      return stbi__errpuc("not PSD", "Corrupt PSD image");
6358+
6359+   /* Check file type version. */
6360+   if (stbi__get16be(s) != 1)
6361+      return stbi__errpuc("wrong version", "Unsupported version of PSD image");
6362+
6363+   /* Skip 6 reserved bytes. */
6364+   stbi__skip(s, 6 );
6365+
6366+   /* Read the number of channels (R, G, B, A, etc). */
6367+   channelCount = stbi__get16be(s);
6368+   if (channelCount < 0 || channelCount > 16)
6369+      return stbi__errpuc("wrong channel count", "Unsupported number of channels in PSD image");
6370+
6371+   /* Read the rows and columns of the image. */
6372+   h = stbi__get32be(s);
6373+   w = stbi__get32be(s);
6374+
6375+   if (h > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
6376+   if (w > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
6377+
6378+   /* Make sure the depth is 8 bits. */
6379+   bitdepth = stbi__get16be(s);
6380+   if (bitdepth != 8 && bitdepth != 16)
6381+      return stbi__errpuc("unsupported bit depth", "PSD bit depth is not 8 or 16 bit");
6382+
6383+   /*
6384+    * Make sure the color mode is RGB.
6385+    * Valid options are:
6386+    *   0: Bitmap
6387+    *   1: Grayscale
6388+    *   2: Indexed color
6389+    *   3: RGB color
6390+    *   4: CMYK color
6391+    *   7: Multichannel
6392+    *   8: Duotone
6393+    *   9: Lab color
6394+    */
6395+   if (stbi__get16be(s) != 3)
6396+      return stbi__errpuc("wrong color format", "PSD is not in RGB color format");
6397+
6398+   /* Skip the Mode Data.  (It's the palette for indexed color; other info for other modes.) */
6399+   stbi__skip(s,stbi__get32be(s) );
6400+
6401+   /* Skip the image resources.  (resolution, pen tool paths, etc) */
6402+   stbi__skip(s, stbi__get32be(s) );
6403+
6404+   /* Skip the reserved data. */
6405+   stbi__skip(s, stbi__get32be(s) );
6406+
6407+   /*
6408+    * Find out if the data is compressed.
6409+    * Known values:
6410+    *   0: no compression
6411+    *   1: RLE compressed
6412+    */
6413+   compression = stbi__get16be(s);
6414+   if (compression > 1)
6415+      return stbi__errpuc("bad compression", "PSD has an unknown compression format");
6416+
6417+   /* Check size */
6418+   if (!stbi__mad3sizes_valid(4, w, h, 0))
6419+      return stbi__errpuc("too large", "Corrupt PSD");
6420+
6421+   /* Create the destination image. */
6422+
6423+   if (!compression && bitdepth == 16 && bpc == 16) {
6424+      out = (stbi_uc *) stbi__malloc_mad3(8, w, h, 0);
6425+      ri->bits_per_channel = 16;
6426+   } else
6427+      out = (stbi_uc *) stbi__malloc(4 * w*h);
6428+
6429+   if (!out) return stbi__errpuc("outofmem", "Out of memory");
6430+   pixelCount = w*h;
6431+
6432+   /*
6433+    * Initialize the data to zero.
6434+    * memset( out, 0, pixelCount * 4 );
6435+    */
6436+
6437+   /* Finally, the image data. */
6438+   if (compression) {
6439+      /*
6440+       * RLE as used by .PSD and .TIFF
6441+       * Loop until you get the number of unpacked bytes you are expecting:
6442+       *     Read the next source byte into n.
6443+       *     If n is between 0 and 127 inclusive, copy the next n+1 bytes literally.
6444+       *     Else if n is between -127 and -1 inclusive, copy the next byte -n+1 times.
6445+       *     Else if n is 128, noop.
6446+       * Endloop
6447+       */
6448+
6449+      /*
6450+       * The RLE-compressed data is preceded by a 2-byte data count for each row in the data,
6451+       * which we're going to just skip.
6452+       */
6453+      stbi__skip(s, h * channelCount * 2 );
6454+
6455+      /* Read the RLE data by channel. */
6456+      for (channel = 0; channel < 4; channel++) {
6457+         stbi_uc *p;
6458+
6459+         p = out+channel;
6460+         if (channel >= channelCount) {
6461+            /* Fill this channel with default data. */
6462+            for (i = 0; i < pixelCount; i++, p += 4)
6463+               *p = (channel == 3 ? 255 : 0);
6464+         } else {
6465+            /* Read the RLE data. */
6466+            if (!stbi__psd_decode_rle(s, p, pixelCount)) {
6467+               STBI_FREE(out);
6468+               return stbi__errpuc("corrupt", "bad RLE data");
6469+            }
6470+         }
6471+      }
6472+
6473+   } else {
6474+      /*
6475+       * We're at the raw image data.  It's each channel in order (Red, Green, Blue, Alpha, ...)
6476+       * where each channel consists of an 8-bit (or 16-bit) value for each pixel in the image.
6477+       */
6478+
6479+      /* Read the data by channel. */
6480+      for (channel = 0; channel < 4; channel++) {
6481+         if (channel >= channelCount) {
6482+            /* Fill this channel with default data. */
6483+            if (bitdepth == 16 && bpc == 16) {
6484+               stbi__uint16 *q = ((stbi__uint16 *) out) + channel;
6485+               stbi__uint16 val = channel == 3 ? 65535 : 0;
6486+               for (i = 0; i < pixelCount; i++, q += 4)
6487+                  *q = val;
6488+            } else {
6489+               stbi_uc *p = out+channel;
6490+               stbi_uc val = channel == 3 ? 255 : 0;
6491+               for (i = 0; i < pixelCount; i++, p += 4)
6492+                  *p = val;
6493+            }
6494+         } else {
6495+            if (ri->bits_per_channel == 16) { /* output bpc */
6496+               stbi__uint16 *q = ((stbi__uint16 *) out) + channel;
6497+               for (i = 0; i < pixelCount; i++, q += 4)
6498+                  *q = (stbi__uint16) stbi__get16be(s);
6499+            } else {
6500+               stbi_uc *p = out+channel;
6501+               if (bitdepth == 16) { /* input bpc */
6502+                  for (i = 0; i < pixelCount; i++, p += 4)
6503+                     *p = (stbi_uc) (stbi__get16be(s) >> 8);
6504+               } else {
6505+                  for (i = 0; i < pixelCount; i++, p += 4)
6506+                     *p = stbi__get8(s);
6507+               }
6508+            }
6509+         }
6510+      }
6511+   }
6512+
6513+   /* remove weird white matte from PSD */
6514+   if (channelCount >= 4) {
6515+      if (ri->bits_per_channel == 16) {
6516+         for (i=0; i < w*h; ++i) {
6517+            stbi__uint16 *pixel = (stbi__uint16 *) out + 4*i;
6518+            if (pixel[3] != 0 && pixel[3] != 65535) {
6519+               float a = pixel[3] / 65535.0f;
6520+               float ra = 1.0f / a;
6521+               float inv_a = 65535.0f * (1 - ra);
6522+               pixel[0] = (stbi__uint16) (pixel[0]*ra + inv_a);
6523+               pixel[1] = (stbi__uint16) (pixel[1]*ra + inv_a);
6524+               pixel[2] = (stbi__uint16) (pixel[2]*ra + inv_a);
6525+            }
6526+         }
6527+      } else {
6528+         for (i=0; i < w*h; ++i) {
6529+            unsigned char *pixel = out + 4*i;
6530+            if (pixel[3] != 0 && pixel[3] != 255) {
6531+               float a = pixel[3] / 255.0f;
6532+               float ra = 1.0f / a;
6533+               float inv_a = 255.0f * (1 - ra);
6534+               pixel[0] = (unsigned char) (pixel[0]*ra + inv_a);
6535+               pixel[1] = (unsigned char) (pixel[1]*ra + inv_a);
6536+               pixel[2] = (unsigned char) (pixel[2]*ra + inv_a);
6537+            }
6538+         }
6539+      }
6540+   }
6541+
6542+   /* convert to desired output format */
6543+   if (req_comp && req_comp != 4) {
6544+      if (ri->bits_per_channel == 16)
6545+         out = (stbi_uc *) stbi__convert_format16((stbi__uint16 *) out, 4, req_comp, w, h);
6546+      else
6547+         out = stbi__convert_format(out, 4, req_comp, w, h);
6548+      if (out == NULL) return out; /* stbi__convert_format frees input on failure */
6549+   }
6550+
6551+   if (comp) *comp = 4;
6552+   *y = h;
6553+   *x = w;
6554+
6555+   return out;
6556+}
6557+#endif
6558+
6559+/***************************************************************************************************/
6560+/*
6561+ * Softimage PIC loader
6562+ * by Tom Seddon
6563+ *
6564+ * See http://softimage.wiki.softimage.com/index.php/INFO:_PIC_file_format
6565+ * See http://ozviz.wasp.uwa.edu.au/~pbourke/dataformats/softimagepic/
6566+ */
6567+
6568+#ifndef STBI_NO_PIC
6569+static int stbi__pic_is4(stbi__context *s,const char *str)
6570+{
6571+   int i;
6572+   for (i=0; i<4; ++i)
6573+      if (stbi__get8(s) != (stbi_uc)str[i])
6574+         return 0;
6575+
6576+   return 1;
6577+}
6578+
6579+static int stbi__pic_test_core(stbi__context *s)
6580+{
6581+   int i;
6582+
6583+   if (!stbi__pic_is4(s,"\x53\x80\xF6\x34"))
6584+      return 0;
6585+
6586+   for(i=0;i<84;++i)
6587+      stbi__get8(s);
6588+
6589+   if (!stbi__pic_is4(s,"PICT"))
6590+      return 0;
6591+
6592+   return 1;
6593+}
6594+
6595+typedef struct
6596+{
6597+   stbi_uc size,type,channel;
6598+} stbi__pic_packet;
6599+
6600+static stbi_uc *stbi__readval(stbi__context *s, int channel, stbi_uc *dest)
6601+{
6602+   int mask=0x80, i;
6603+
6604+   for (i=0; i<4; ++i, mask>>=1) {
6605+      if (channel & mask) {
6606+         if (stbi__at_eof(s)) return stbi__errpuc("bad file","PIC file too short");
6607+         dest[i]=stbi__get8(s);
6608+      }
6609+   }
6610+
6611+   return dest;
6612+}
6613+
6614+static void stbi__copyval(int channel,stbi_uc *dest,const stbi_uc *src)
6615+{
6616+   int mask=0x80,i;
6617+
6618+   for (i=0;i<4; ++i, mask>>=1)
6619+      if (channel&mask)
6620+         dest[i]=src[i];
6621+}
6622+
6623+static stbi_uc *stbi__pic_load_core(stbi__context *s,int width,int height,int *comp, stbi_uc *result)
6624+{
6625+   int act_comp=0,num_packets=0,y,chained;
6626+   stbi__pic_packet packets[10];
6627+
6628+   /*
6629+    * this will (should...) cater for even some bizarre stuff like having data
6630+    * for the same channel in multiple packets.
6631+    */
6632+   do {
6633+      stbi__pic_packet *packet;
6634+
6635+      if (num_packets==sizeof(packets)/sizeof(packets[0]))
6636+         return stbi__errpuc("bad format","too many packets");
6637+
6638+      packet = &packets[num_packets++];
6639+
6640+      chained = stbi__get8(s);
6641+      packet->size    = stbi__get8(s);
6642+      packet->type    = stbi__get8(s);
6643+      packet->channel = stbi__get8(s);
6644+
6645+      act_comp |= packet->channel;
6646+
6647+      if (stbi__at_eof(s))          return stbi__errpuc("bad file","file too short (reading packets)");
6648+      if (packet->size != 8)  return stbi__errpuc("bad format","packet isn't 8bpp");
6649+   } while (chained);
6650+
6651+   *comp = (act_comp & 0x10 ? 4 : 3); /* has alpha channel? */
6652+
6653+   for(y=0; y<height; ++y) {
6654+      int packet_idx;
6655+
6656+      for(packet_idx=0; packet_idx < num_packets; ++packet_idx) {
6657+         stbi__pic_packet *packet = &packets[packet_idx];
6658+         stbi_uc *dest = result+y*width*4;
6659+
6660+         switch (packet->type) {
6661+            default:
6662+               return stbi__errpuc("bad format","packet has bad compression type");
6663+
6664+            case 0: { /* uncompressed */
6665+               int x;
6666+
6667+               for(x=0;x<width;++x, dest+=4)
6668+                  if (!stbi__readval(s,packet->channel,dest))
6669+                     return 0;
6670+               break;
6671+            }
6672+
6673+            case 1: /* Pure RLE */
6674+               {
6675+                  int left=width, i;
6676+
6677+                  while (left>0) {
6678+                     stbi_uc count,value[4];
6679+
6680+                     count=stbi__get8(s);
6681+                     if (stbi__at_eof(s))   return stbi__errpuc("bad file","file too short (pure read count)");
6682+
6683+                     if (count > left)
6684+                        count = (stbi_uc) left;
6685+
6686+                     if (!stbi__readval(s,packet->channel,value))  return 0;
6687+
6688+                     for(i=0; i<count; ++i,dest+=4)
6689+                        stbi__copyval(packet->channel,dest,value);
6690+                     left -= count;
6691+                  }
6692+               }
6693+               break;
6694+
6695+            case 2: { /* Mixed RLE */
6696+               int left=width;
6697+               while (left>0) {
6698+                  int count = stbi__get8(s), i;
6699+                  if (stbi__at_eof(s))  return stbi__errpuc("bad file","file too short (mixed read count)");
6700+
6701+                  if (count >= 128) { /* Repeated */
6702+                     stbi_uc value[4];
6703+
6704+                     if (count==128)
6705+                        count = stbi__get16be(s);
6706+                     else
6707+                        count -= 127;
6708+                     if (count > left)
6709+                        return stbi__errpuc("bad file","scanline overrun");
6710+
6711+                     if (!stbi__readval(s,packet->channel,value))
6712+                        return 0;
6713+
6714+                     for(i=0;i<count;++i, dest += 4)
6715+                        stbi__copyval(packet->channel,dest,value);
6716+                  } else { /* Raw */
6717+                     ++count;
6718+                     if (count>left) return stbi__errpuc("bad file","scanline overrun");
6719+
6720+                     for(i=0;i<count;++i, dest+=4)
6721+                        if (!stbi__readval(s,packet->channel,dest))
6722+                           return 0;
6723+                  }
6724+                  left-=count;
6725+               }
6726+               break;
6727+            }
6728+         }
6729+      }
6730+   }
6731+
6732+   return result;
6733+}
6734+
6735+static void *stbi__pic_load(stbi__context *s,int *px,int *py,int *comp,int req_comp, stbi__result_info *ri)
6736+{
6737+   stbi_uc *result;
6738+   int i, x,y, internal_comp;
6739+   STBI_NOTUSED(ri);
6740+
6741+   if (!comp) comp = &internal_comp;
6742+
6743+   for (i=0; i<92; ++i)
6744+      stbi__get8(s);
6745+
6746+   x = stbi__get16be(s);
6747+   y = stbi__get16be(s);
6748+
6749+   if (y > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
6750+   if (x > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
6751+
6752+   if (stbi__at_eof(s))  return stbi__errpuc("bad file","file too short (pic header)");
6753+   if (!stbi__mad3sizes_valid(x, y, 4, 0)) return stbi__errpuc("too large", "PIC image too large to decode");
6754+
6755+   stbi__get32be(s); /* skip `ratio' */
6756+   stbi__get16be(s); /* skip `fields' */
6757+   stbi__get16be(s); /* skip `pad' */
6758+
6759+   /* intermediate buffer is RGBA */
6760+   result = (stbi_uc *) stbi__malloc_mad3(x, y, 4, 0);
6761+   if (!result) return stbi__errpuc("outofmem", "Out of memory");
6762+   memset(result, 0xff, x*y*4);
6763+
6764+   if (!stbi__pic_load_core(s,x,y,comp, result)) {
6765+      STBI_FREE(result);
6766+      result=0;
6767+   }
6768+   *px = x;
6769+   *py = y;
6770+   if (req_comp == 0) req_comp = *comp;
6771+   result=stbi__convert_format(result,4,req_comp,x,y);
6772+
6773+   return result;
6774+}
6775+
6776+static int stbi__pic_test(stbi__context *s)
6777+{
6778+   int r = stbi__pic_test_core(s);
6779+   stbi__rewind(s);
6780+   return r;
6781+}
6782+#endif
6783+
6784+/***************************************************************************************************/
6785+/* GIF loader -- public domain by Jean-Marc Lienher -- simplified/shrunk by stb */
6786+
6787+#ifndef STBI_NO_GIF
6788+typedef struct
6789+{
6790+   stbi__int16 prefix;
6791+   stbi_uc first;
6792+   stbi_uc suffix;
6793+} stbi__gif_lzw;
6794+
6795+typedef struct
6796+{
6797+   int w,h;
6798+   stbi_uc *out; /* output buffer (always 4 components) */
6799+   stbi_uc *background; /* The current "background" as far as a gif is concerned */
6800+   stbi_uc *history;
6801+   int flags, bgindex, ratio, transparent, eflags;
6802+   stbi_uc  pal[256][4];
6803+   stbi_uc lpal[256][4];
6804+   stbi__gif_lzw codes[8192];
6805+   stbi_uc *color_table;
6806+   int parse, step;
6807+   int lflags;
6808+   int start_x, start_y;
6809+   int max_x, max_y;
6810+   int cur_x, cur_y;
6811+   int line_size;
6812+   int delay;
6813+} stbi__gif;
6814+
6815+static int stbi__gif_test_raw(stbi__context *s)
6816+{
6817+   int sz;
6818+   if (stbi__get8(s) != 'G' || stbi__get8(s) != 'I' || stbi__get8(s) != 'F' || stbi__get8(s) != '8') return 0;
6819+   sz = stbi__get8(s);
6820+   if (sz != '9' && sz != '7') return 0;
6821+   if (stbi__get8(s) != 'a') return 0;
6822+   return 1;
6823+}
6824+
6825+static int stbi__gif_test(stbi__context *s)
6826+{
6827+   int r = stbi__gif_test_raw(s);
6828+   stbi__rewind(s);
6829+   return r;
6830+}
6831+
6832+static void stbi__gif_parse_colortable(stbi__context *s, stbi_uc pal[256][4], int num_entries, int transp)
6833+{
6834+   int i;
6835+   for (i=0; i < num_entries; ++i) {
6836+      pal[i][2] = stbi__get8(s);
6837+      pal[i][1] = stbi__get8(s);
6838+      pal[i][0] = stbi__get8(s);
6839+      pal[i][3] = transp == i ? 0 : 255;
6840+   }
6841+}
6842+
6843+static int stbi__gif_header(stbi__context *s, stbi__gif *g, int *comp, int is_info)
6844+{
6845+   stbi_uc version;
6846+   if (stbi__get8(s) != 'G' || stbi__get8(s) != 'I' || stbi__get8(s) != 'F' || stbi__get8(s) != '8')
6847+      return stbi__err("not GIF", "Corrupt GIF");
6848+
6849+   version = stbi__get8(s);
6850+   if (version != '7' && version != '9')    return stbi__err("not GIF", "Corrupt GIF");
6851+   if (stbi__get8(s) != 'a')                return stbi__err("not GIF", "Corrupt GIF");
6852+
6853+   stbi__g_failure_reason = "";
6854+   g->w = stbi__get16le(s);
6855+   g->h = stbi__get16le(s);
6856+   g->flags = stbi__get8(s);
6857+   g->bgindex = stbi__get8(s);
6858+   g->ratio = stbi__get8(s);
6859+   g->transparent = -1;
6860+
6861+   if (g->w > STBI_MAX_DIMENSIONS) return stbi__err("too large","Very large image (corrupt?)");
6862+   if (g->h > STBI_MAX_DIMENSIONS) return stbi__err("too large","Very large image (corrupt?)");
6863+
6864+   if (comp != 0) *comp = 4; /* can't actually tell whether it's 3 or 4 until we parse the comments */
6865+
6866+   if (is_info) return 1;
6867+
6868+   if (g->flags & 0x80)
6869+      stbi__gif_parse_colortable(s,g->pal, 2 << (g->flags & 7), -1);
6870+
6871+   return 1;
6872+}
6873+
6874+static int stbi__gif_info_raw(stbi__context *s, int *x, int *y, int *comp)
6875+{
6876+   stbi__gif* g = (stbi__gif*) stbi__malloc(sizeof(stbi__gif));
6877+   if (!g) return stbi__err("outofmem", "Out of memory");
6878+   if (!stbi__gif_header(s, g, comp, 1)) {
6879+      STBI_FREE(g);
6880+      stbi__rewind( s );
6881+      return 0;
6882+   }
6883+   if (x) *x = g->w;
6884+   if (y) *y = g->h;
6885+   STBI_FREE(g);
6886+   return 1;
6887+}
6888+
6889+static void stbi__out_gif_code(stbi__gif *g, stbi__uint16 code)
6890+{
6891+   stbi_uc *p, *c;
6892+   int idx;
6893+
6894+   /*
6895+    * recurse to decode the prefixes, since the linked-list is backwards,
6896+    * and working backwards through an interleaved image would be nasty
6897+    */
6898+   if (g->codes[code].prefix >= 0)
6899+      stbi__out_gif_code(g, g->codes[code].prefix);
6900+
6901+   if (g->cur_y >= g->max_y) return;
6902+
6903+   idx = g->cur_x + g->cur_y;
6904+   p = &g->out[idx];
6905+   g->history[idx / 4] = 1;
6906+
6907+   c = &g->color_table[g->codes[code].suffix * 4];
6908+   if (c[3] > 128) { /* don't render transparent pixels; */
6909+      p[0] = c[2];
6910+      p[1] = c[1];
6911+      p[2] = c[0];
6912+      p[3] = c[3];
6913+   }
6914+   g->cur_x += 4;
6915+
6916+   if (g->cur_x >= g->max_x) {
6917+      g->cur_x = g->start_x;
6918+      g->cur_y += g->step;
6919+
6920+      while (g->cur_y >= g->max_y && g->parse > 0) {
6921+         g->step = (1 << g->parse) * g->line_size;
6922+         g->cur_y = g->start_y + (g->step >> 1);
6923+         --g->parse;
6924+      }
6925+   }
6926+}
6927+
6928+static stbi_uc *stbi__process_gif_raster(stbi__context *s, stbi__gif *g)
6929+{
6930+   stbi_uc lzw_cs;
6931+   stbi__int32 len, init_code;
6932+   stbi__uint32 first;
6933+   stbi__int32 codesize, codemask, avail, oldcode, bits, valid_bits, clear;
6934+   stbi__gif_lzw *p;
6935+
6936+   lzw_cs = stbi__get8(s);
6937+   if (lzw_cs > 12) return NULL;
6938+   clear = 1 << lzw_cs;
6939+   first = 1;
6940+   codesize = lzw_cs + 1;
6941+   codemask = (1 << codesize) - 1;
6942+   bits = 0;
6943+   valid_bits = 0;
6944+   for (init_code = 0; init_code < clear; init_code++) {
6945+      g->codes[init_code].prefix = -1;
6946+      g->codes[init_code].first = (stbi_uc) init_code;
6947+      g->codes[init_code].suffix = (stbi_uc) init_code;
6948+   }
6949+
6950+   /* support no starting clear code */
6951+   avail = clear+2;
6952+   oldcode = -1;
6953+
6954+   len = 0;
6955+   for(;;) {
6956+      if (valid_bits < codesize) {
6957+         if (len == 0) {
6958+            len = stbi__get8(s); /* start new block */
6959+            if (len == 0)
6960+               return g->out;
6961+         }
6962+         --len;
6963+         bits |= (stbi__int32) stbi__get8(s) << valid_bits;
6964+         valid_bits += 8;
6965+      } else {
6966+         stbi__int32 code = bits & codemask;
6967+         bits >>= codesize;
6968+         valid_bits -= codesize;
6969+         /* @OPTIMIZE: is there some way we can accelerate the non-clear path? */
6970+         if (code == clear) { /* clear code */
6971+            codesize = lzw_cs + 1;
6972+            codemask = (1 << codesize) - 1;
6973+            avail = clear + 2;
6974+            oldcode = -1;
6975+            first = 0;
6976+         } else if (code == clear + 1) { /* end of stream code */
6977+            stbi__skip(s, len);
6978+            while ((len = stbi__get8(s)) > 0)
6979+               stbi__skip(s,len);
6980+            return g->out;
6981+         } else if (code <= avail) {
6982+            if (first) {
6983+               return stbi__errpuc("no clear code", "Corrupt GIF");
6984+            }
6985+
6986+            if (oldcode >= 0) {
6987+               p = &g->codes[avail++];
6988+               if (avail > 8192) {
6989+                  return stbi__errpuc("too many codes", "Corrupt GIF");
6990+               }
6991+
6992+               p->prefix = (stbi__int16) oldcode;
6993+               p->first = g->codes[oldcode].first;
6994+               p->suffix = (code == avail) ? p->first : g->codes[code].first;
6995+            } else if (code == avail)
6996+               return stbi__errpuc("illegal code in raster", "Corrupt GIF");
6997+
6998+            stbi__out_gif_code(g, (stbi__uint16) code);
6999+
7000+            if ((avail & codemask) == 0 && avail <= 0x0FFF) {
7001+               codesize++;
7002+               codemask = (1 << codesize) - 1;
7003+            }
7004+
7005+            oldcode = code;
7006+         } else {
7007+            return stbi__errpuc("illegal code in raster", "Corrupt GIF");
7008+         }
7009+      }
7010+   }
7011+}
7012+
7013+/*
7014+ * this function is designed to support animated gifs, although stb_image doesn't support it
7015+ * two back is the image from two frames ago, used for a very specific disposal format
7016+ */
7017+static stbi_uc *stbi__gif_load_next(stbi__context *s, stbi__gif *g, int *comp, int req_comp, stbi_uc *two_back)
7018+{
7019+   int dispose;
7020+   int first_frame;
7021+   int pi;
7022+   int pcount;
7023+   STBI_NOTUSED(req_comp);
7024+
7025+   /* on first frame, any non-written pixels get the background colour (non-transparent) */
7026+   first_frame = 0;
7027+   if (g->out == 0) {
7028+      if (!stbi__gif_header(s, g, comp,0)) return 0; /* stbi__g_failure_reason set by stbi__gif_header */
7029+      if (!stbi__mad3sizes_valid(4, g->w, g->h, 0))
7030+         return stbi__errpuc("too large", "GIF image is too large");
7031+      pcount = g->w * g->h;
7032+      g->out = (stbi_uc *) stbi__malloc(4 * pcount);
7033+      g->background = (stbi_uc *) stbi__malloc(4 * pcount);
7034+      g->history = (stbi_uc *) stbi__malloc(pcount);
7035+      if (!g->out || !g->background || !g->history)
7036+         return stbi__errpuc("outofmem", "Out of memory");
7037+
7038+      /*
7039+       * image is treated as "transparent" at the start - ie, nothing overwrites the current background;
7040+       * background colour is only used for pixels that are not rendered first frame, after that "background"
7041+       * color refers to the color that was there the previous frame.
7042+       */
7043+      memset(g->out, 0x00, 4 * pcount);
7044+      memset(g->background, 0x00, 4 * pcount); /* state of the background (starts transparent) */
7045+      memset(g->history, 0x00, pcount); /* pixels that were affected previous frame */
7046+      first_frame = 1;
7047+   } else {
7048+      /* second frame - how do we dispose of the previous one? */
7049+      dispose = (g->eflags & 0x1C) >> 2;
7050+      pcount = g->w * g->h;
7051+
7052+      if ((dispose == 3) && (two_back == 0)) {
7053+         dispose = 2; /* if I don't have an image to revert back to, default to the old background */
7054+      }
7055+
7056+      if (dispose == 3) { /* use previous graphic */
7057+         for (pi = 0; pi < pcount; ++pi) {
7058+            if (g->history[pi]) {
7059+               memcpy( &g->out[pi * 4], &two_back[pi * 4], 4 );
7060+            }
7061+         }
7062+      } else if (dispose == 2) {
7063+         /* restore what was changed last frame to background before that frame; */
7064+         for (pi = 0; pi < pcount; ++pi) {
7065+            if (g->history[pi]) {
7066+               memcpy( &g->out[pi * 4], &g->background[pi * 4], 4 );
7067+            }
7068+         }
7069+      } else {
7070+         /*
7071+          * This is a non-disposal case eithe way, so just
7072+          * leave the pixels as is, and they will become the new background
7073+          * 1: do not dispose
7074+          * 0:  not specified.
7075+          */
7076+      }
7077+
7078+      /* background is what out is after the undoing of the previou frame; */
7079+      memcpy( g->background, g->out, 4 * g->w * g->h );
7080+   }
7081+
7082+   /* clear my history; */
7083+   memset( g->history, 0x00, g->w * g->h ); /* pixels that were affected previous frame */
7084+
7085+   for (;;) {
7086+      int tag = stbi__get8(s);
7087+      switch (tag) {
7088+         case 0x2C: /* Image Descriptor */
7089+         {
7090+            stbi__int32 x, y, w, h;
7091+            stbi_uc *o;
7092+
7093+            x = stbi__get16le(s);
7094+            y = stbi__get16le(s);
7095+            w = stbi__get16le(s);
7096+            h = stbi__get16le(s);
7097+            if (((x + w) > (g->w)) || ((y + h) > (g->h)))
7098+               return stbi__errpuc("bad Image Descriptor", "Corrupt GIF");
7099+
7100+            g->line_size = g->w * 4;
7101+            g->start_x = x * 4;
7102+            g->start_y = y * g->line_size;
7103+            g->max_x   = g->start_x + w * 4;
7104+            g->max_y   = g->start_y + h * g->line_size;
7105+            g->cur_x   = g->start_x;
7106+            g->cur_y   = g->start_y;
7107+
7108+            /*
7109+             * if the width of the specified rectangle is 0, that means
7110+             * we may not see *any* pixels or the image is malformed;
7111+             * to make sure this is caught, move the current y down to
7112+             * max_y (which is what out_gif_code checks).
7113+             */
7114+            if (w == 0)
7115+               g->cur_y = g->max_y;
7116+
7117+            g->lflags = stbi__get8(s);
7118+
7119+            if (g->lflags & 0x40) {
7120+               g->step = 8 * g->line_size; /* first interlaced spacing */
7121+               g->parse = 3;
7122+            } else {
7123+               g->step = g->line_size;
7124+               g->parse = 0;
7125+            }
7126+
7127+            if (g->lflags & 0x80) {
7128+               stbi__gif_parse_colortable(s,g->lpal, 2 << (g->lflags & 7), g->eflags & 0x01 ? g->transparent : -1);
7129+               g->color_table = (stbi_uc *) g->lpal;
7130+            } else if (g->flags & 0x80) {
7131+               g->color_table = (stbi_uc *) g->pal;
7132+            } else
7133+               return stbi__errpuc("missing color table", "Corrupt GIF");
7134+
7135+            o = stbi__process_gif_raster(s, g);
7136+            if (!o) return NULL;
7137+
7138+            /* if this was the first frame, */
7139+            pcount = g->w * g->h;
7140+            if (first_frame && (g->bgindex > 0)) {
7141+               /* if first frame, any pixel not drawn to gets the background color */
7142+               for (pi = 0; pi < pcount; ++pi) {
7143+                  if (g->history[pi] == 0) {
7144+                     g->pal[g->bgindex][3] = 255; /* just in case it was made transparent, undo that; It will be reset next frame if need be; */
7145+                     memcpy( &g->out[pi * 4], &g->pal[g->bgindex], 4 );
7146+                  }
7147+               }
7148+            }
7149+
7150+            return o;
7151+         }
7152+
7153+         case 0x21: /* Comment Extension. */
7154+         {
7155+            int len;
7156+            int ext = stbi__get8(s);
7157+            if (ext == 0xF9) { /* Graphic Control Extension. */
7158+               len = stbi__get8(s);
7159+               if (len == 4) {
7160+                  g->eflags = stbi__get8(s);
7161+                  g->delay = 10 * stbi__get16le(s); /* delay - 1/100th of a second, saving as 1/1000ths. */
7162+
7163+                  /* unset old transparent */
7164+                  if (g->transparent >= 0) {
7165+                     g->pal[g->transparent][3] = 255;
7166+                  }
7167+                  if (g->eflags & 0x01) {
7168+                     g->transparent = stbi__get8(s);
7169+                     if (g->transparent >= 0) {
7170+                        g->pal[g->transparent][3] = 0;
7171+                     }
7172+                  } else {
7173+                     /* don't need transparent */
7174+                     stbi__skip(s, 1);
7175+                     g->transparent = -1;
7176+                  }
7177+               } else {
7178+                  stbi__skip(s, len);
7179+                  break;
7180+               }
7181+            }
7182+            while ((len = stbi__get8(s)) != 0) {
7183+               stbi__skip(s, len);
7184+            }
7185+            break;
7186+         }
7187+
7188+         case 0x3B: /* gif stream termination code */
7189+            return (stbi_uc *) s; /* using '1' causes warning on some compilers */
7190+
7191+         default:
7192+            return stbi__errpuc("unknown code", "Corrupt GIF");
7193+      }
7194+   }
7195+}
7196+
7197+static void *stbi__load_gif_main_outofmem(stbi__gif *g, stbi_uc *out, int **delays)
7198+{
7199+   STBI_FREE(g->out);
7200+   STBI_FREE(g->history);
7201+   STBI_FREE(g->background);
7202+
7203+   if (out) STBI_FREE(out);
7204+   if (delays && *delays) STBI_FREE(*delays);
7205+   return stbi__errpuc("outofmem", "Out of memory");
7206+}
7207+
7208+static void *stbi__load_gif_main(stbi__context *s, int **delays, int *x, int *y, int *z, int *comp, int req_comp)
7209+{
7210+   if (stbi__gif_test(s)) {
7211+      int layers = 0;
7212+      stbi_uc *u = 0;
7213+      stbi_uc *out = 0;
7214+      stbi_uc *two_back = 0;
7215+      stbi__gif g;
7216+      int stride;
7217+      int out_size = 0;
7218+      int delays_size = 0;
7219+
7220+      STBI_NOTUSED(out_size);
7221+      STBI_NOTUSED(delays_size);
7222+
7223+      memset(&g, 0, sizeof(g));
7224+      if (delays) {
7225+         *delays = 0;
7226+      }
7227+
7228+      do {
7229+         u = stbi__gif_load_next(s, &g, comp, req_comp, two_back);
7230+         if (u == (stbi_uc *) s) u = 0; /* end of animated gif marker */
7231+
7232+         if (u) {
7233+            *x = g.w;
7234+            *y = g.h;
7235+            ++layers;
7236+            stride = g.w * g.h * 4;
7237+
7238+            if (out) {
7239+               void *tmp = (stbi_uc*) STBI_REALLOC_SIZED( out, out_size, layers * stride );
7240+               if (!tmp)
7241+                  return stbi__load_gif_main_outofmem(&g, out, delays);
7242+               else {
7243+                   out = (stbi_uc*) tmp;
7244+                   out_size = layers * stride;
7245+               }
7246+
7247+               if (delays) {
7248+                  int *new_delays = (int*) STBI_REALLOC_SIZED( *delays, delays_size, sizeof(int) * layers );
7249+                  if (!new_delays)
7250+                     return stbi__load_gif_main_outofmem(&g, out, delays);
7251+                  *delays = new_delays;
7252+                  delays_size = layers * sizeof(int);
7253+               }
7254+            } else {
7255+               out = (stbi_uc*)stbi__malloc( layers * stride );
7256+               if (!out)
7257+                  return stbi__load_gif_main_outofmem(&g, out, delays);
7258+               out_size = layers * stride;
7259+               if (delays) {
7260+                  *delays = (int*) stbi__malloc( layers * sizeof(int) );
7261+                  if (!*delays)
7262+                     return stbi__load_gif_main_outofmem(&g, out, delays);
7263+                  delays_size = layers * sizeof(int);
7264+               }
7265+            }
7266+            memcpy( out + ((layers - 1) * stride), u, stride );
7267+            if (layers >= 2) {
7268+               two_back = out - 2 * stride;
7269+            }
7270+
7271+            if (delays) {
7272+               (*delays)[layers - 1U] = g.delay;
7273+            }
7274+         }
7275+      } while (u != 0);
7276+
7277+      /* free temp buffer; */
7278+      STBI_FREE(g.out);
7279+      STBI_FREE(g.history);
7280+      STBI_FREE(g.background);
7281+
7282+      /* do the final conversion after loading everything; */
7283+      if (req_comp && req_comp != 4)
7284+         out = stbi__convert_format(out, 4, req_comp, layers * g.w, g.h);
7285+
7286+      *z = layers;
7287+      return out;
7288+   } else {
7289+      return stbi__errpuc("not GIF", "Image was not as a gif type.");
7290+   }
7291+}
7292+
7293+static void *stbi__gif_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri)
7294+{
7295+   stbi_uc *u = 0;
7296+   stbi__gif g;
7297+   memset(&g, 0, sizeof(g));
7298+   STBI_NOTUSED(ri);
7299+
7300+   u = stbi__gif_load_next(s, &g, comp, req_comp, 0);
7301+   if (u == (stbi_uc *) s) u = 0; /* end of animated gif marker */
7302+   if (u) {
7303+      *x = g.w;
7304+      *y = g.h;
7305+
7306+      /*
7307+       * moved conversion to after successful load so that the same
7308+       * can be done for multiple frames.
7309+       */
7310+      if (req_comp && req_comp != 4)
7311+         u = stbi__convert_format(u, 4, req_comp, g.w, g.h);
7312+   } else if (g.out) {
7313+      /* if there was an error and we allocated an image buffer, free it! */
7314+      STBI_FREE(g.out);
7315+   }
7316+
7317+   /* free buffers needed for multiple frame loading; */
7318+   STBI_FREE(g.history);
7319+   STBI_FREE(g.background);
7320+
7321+   return u;
7322+}
7323+
7324+static int stbi__gif_info(stbi__context *s, int *x, int *y, int *comp)
7325+{
7326+   return stbi__gif_info_raw(s,x,y,comp);
7327+}
7328+#endif
7329+
7330+/***************************************************************************************************/
7331+/*
7332+ * Radiance RGBE HDR loader
7333+ * originally by Nicolas Schulz
7334+ */
7335+#ifndef STBI_NO_HDR
7336+static int stbi__hdr_test_core(stbi__context *s, const char *signature)
7337+{
7338+   int i;
7339+   for (i=0; signature[i]; ++i)
7340+      if (stbi__get8(s) != signature[i])
7341+          return 0;
7342+   stbi__rewind(s);
7343+   return 1;
7344+}
7345+
7346+static int stbi__hdr_test(stbi__context* s)
7347+{
7348+   int r = stbi__hdr_test_core(s, "#?RADIANCE\n");
7349+   stbi__rewind(s);
7350+   if(!r) {
7351+       r = stbi__hdr_test_core(s, "#?RGBE\n");
7352+       stbi__rewind(s);
7353+   }
7354+   return r;
7355+}
7356+
7357+#define STBI__HDR_BUFLEN  1024
7358+static char *stbi__hdr_gettoken(stbi__context *z, char *buffer)
7359+{
7360+   int len=0;
7361+   char c = '\0';
7362+
7363+   c = (char) stbi__get8(z);
7364+
7365+   while (!stbi__at_eof(z) && c != '\n') {
7366+      buffer[len++] = c;
7367+      if (len == STBI__HDR_BUFLEN-1) {
7368+         /* flush to end of line */
7369+         while (!stbi__at_eof(z) && stbi__get8(z) != '\n')
7370+            ;
7371+         break;
7372+      }
7373+      c = (char) stbi__get8(z);
7374+   }
7375+
7376+   buffer[len] = 0;
7377+   return buffer;
7378+}
7379+
7380+static void stbi__hdr_convert(float *output, stbi_uc *input, int req_comp)
7381+{
7382+   if ( input[3] != 0 ) {
7383+      float f1;
7384+      /* Exponent */
7385+      f1 = (float) ldexp(1.0f, input[3] - (int)(128 + 8));
7386+      if (req_comp <= 2)
7387+         output[0] = (input[0] + input[1] + input[2]) * f1 / 3;
7388+      else {
7389+         output[0] = input[0] * f1;
7390+         output[1] = input[1] * f1;
7391+         output[2] = input[2] * f1;
7392+      }
7393+      if (req_comp == 2) output[1] = 1;
7394+      if (req_comp == 4) output[3] = 1;
7395+   } else {
7396+      switch (req_comp) {
7397+         case 4: output[3] = 1; /* fallthrough */
7398+         case 3: output[0] = output[1] = output[2] = 0;
7399+                 break;
7400+         case 2: output[1] = 1; /* fallthrough */
7401+         case 1: output[0] = 0;
7402+                 break;
7403+      }
7404+   }
7405+}
7406+
7407+static float *stbi__hdr_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri)
7408+{
7409+   char buffer[STBI__HDR_BUFLEN];
7410+   char *token;
7411+   int valid = 0;
7412+   int width, height;
7413+   stbi_uc *scanline;
7414+   float *hdr_data;
7415+   int len;
7416+   unsigned char count, value;
7417+   int i, j, k, c1,c2, z;
7418+   const char *headerToken;
7419+   STBI_NOTUSED(ri);
7420+
7421+   /* Check identifier */
7422+   headerToken = stbi__hdr_gettoken(s,buffer);
7423+   if (strcmp(headerToken, "#?RADIANCE") != 0 && strcmp(headerToken, "#?RGBE") != 0)
7424+      return stbi__errpf("not HDR", "Corrupt HDR image");
7425+
7426+   /* Parse header */
7427+   for(;;) {
7428+      token = stbi__hdr_gettoken(s,buffer);
7429+      if (token[0] == 0) break;
7430+      if (strcmp(token, "FORMAT=32-bit_rle_rgbe") == 0) valid = 1;
7431+   }
7432+
7433+   if (!valid)    return stbi__errpf("unsupported format", "Unsupported HDR format");
7434+
7435+   /*
7436+    * Parse width and height
7437+    * can't use sscanf() if we're not using stdio!
7438+    */
7439+   token = stbi__hdr_gettoken(s,buffer);
7440+   if (strncmp(token, "-Y ", 3))  return stbi__errpf("unsupported data layout", "Unsupported HDR format");
7441+   token += 3;
7442+   height = (int) strtol(token, &token, 10);
7443+   while (*token == ' ') ++token;
7444+   if (strncmp(token, "+X ", 3))  return stbi__errpf("unsupported data layout", "Unsupported HDR format");
7445+   token += 3;
7446+   width = (int) strtol(token, NULL, 10);
7447+
7448+   if (height > STBI_MAX_DIMENSIONS) return stbi__errpf("too large","Very large image (corrupt?)");
7449+   if (width > STBI_MAX_DIMENSIONS) return stbi__errpf("too large","Very large image (corrupt?)");
7450+
7451+   *x = width;
7452+   *y = height;
7453+
7454+   if (comp) *comp = 3;
7455+   if (req_comp == 0) req_comp = 3;
7456+
7457+   if (!stbi__mad4sizes_valid(width, height, req_comp, sizeof(float), 0))
7458+      return stbi__errpf("too large", "HDR image is too large");
7459+
7460+   /* Read data */
7461+   hdr_data = (float *) stbi__malloc_mad4(width, height, req_comp, sizeof(float), 0);
7462+   if (!hdr_data)
7463+      return stbi__errpf("outofmem", "Out of memory");
7464+
7465+   /*
7466+    * Load image data
7467+    * image data is stored as some number of sca
7468+    */
7469+   if ( width < 8 || width >= 32768) {
7470+      /* Read flat data */
7471+      for (j=0; j < height; ++j) {
7472+         for (i=0; i < width; ++i) {
7473+            stbi_uc rgbe[4];
7474+           main_decode_loop:
7475+            stbi__getn(s, rgbe, 4);
7476+            stbi__hdr_convert(hdr_data + j * width * req_comp + i * req_comp, rgbe, req_comp);
7477+         }
7478+      }
7479+   } else {
7480+      /* Read RLE-encoded data */
7481+      scanline = NULL;
7482+
7483+      for (j = 0; j < height; ++j) {
7484+         c1 = stbi__get8(s);
7485+         c2 = stbi__get8(s);
7486+         len = stbi__get8(s);
7487+         if (c1 != 2 || c2 != 2 || (len & 0x80)) {
7488+            /*
7489+             * not run-length encoded, so we have to actually use THIS data as a decoded
7490+             * pixel (note this can't be a valid pixel--one of RGB must be >= 128)
7491+             */
7492+            stbi_uc rgbe[4];
7493+            rgbe[0] = (stbi_uc) c1;
7494+            rgbe[1] = (stbi_uc) c2;
7495+            rgbe[2] = (stbi_uc) len;
7496+            rgbe[3] = (stbi_uc) stbi__get8(s);
7497+            stbi__hdr_convert(hdr_data, rgbe, req_comp);
7498+            i = 1;
7499+            j = 0;
7500+            STBI_FREE(scanline);
7501+            goto main_decode_loop; /* yes, this makes no sense */
7502+         }
7503+         len <<= 8;
7504+         len |= stbi__get8(s);
7505+         if (len != width) { STBI_FREE(hdr_data); STBI_FREE(scanline); return stbi__errpf("invalid decoded scanline length", "corrupt HDR"); }
7506+         if (scanline == NULL) {
7507+            scanline = (stbi_uc *) stbi__malloc_mad2(width, 4, 0);
7508+            if (!scanline) {
7509+               STBI_FREE(hdr_data);
7510+               return stbi__errpf("outofmem", "Out of memory");
7511+            }
7512+         }
7513+
7514+         for (k = 0; k < 4; ++k) {
7515+            int nleft;
7516+            i = 0;
7517+            while ((nleft = width - i) > 0) {
7518+               count = stbi__get8(s);
7519+               if (count > 128) {
7520+                  /* Run */
7521+                  value = stbi__get8(s);
7522+                  count -= 128;
7523+                  if ((count == 0) || (count > nleft)) { STBI_FREE(hdr_data); STBI_FREE(scanline); return stbi__errpf("corrupt", "bad RLE data in HDR"); }
7524+                  for (z = 0; z < count; ++z)
7525+                     scanline[i++ * 4 + k] = value;
7526+               } else {
7527+                  /* Dump */
7528+                  if ((count == 0) || (count > nleft)) { STBI_FREE(hdr_data); STBI_FREE(scanline); return stbi__errpf("corrupt", "bad RLE data in HDR"); }
7529+                  for (z = 0; z < count; ++z)
7530+                     scanline[i++ * 4 + k] = stbi__get8(s);
7531+               }
7532+            }
7533+         }
7534+         for (i=0; i < width; ++i)
7535+            stbi__hdr_convert(hdr_data+(j*width + i)*req_comp, scanline + i*4, req_comp);
7536+      }
7537+      if (scanline)
7538+         STBI_FREE(scanline);
7539+   }
7540+
7541+   return hdr_data;
7542+}
7543+
7544+static int stbi__hdr_info(stbi__context *s, int *x, int *y, int *comp)
7545+{
7546+   char buffer[STBI__HDR_BUFLEN];
7547+   char *token;
7548+   int valid = 0;
7549+   int dummy;
7550+
7551+   if (!x) x = &dummy;
7552+   if (!y) y = &dummy;
7553+   if (!comp) comp = &dummy;
7554+
7555+   if (stbi__hdr_test(s) == 0) {
7556+       stbi__rewind( s );
7557+       return 0;
7558+   }
7559+
7560+   for(;;) {
7561+      token = stbi__hdr_gettoken(s,buffer);
7562+      if (token[0] == 0) break;
7563+      if (strcmp(token, "FORMAT=32-bit_rle_rgbe") == 0) valid = 1;
7564+   }
7565+
7566+   if (!valid) {
7567+       stbi__rewind( s );
7568+       return 0;
7569+   }
7570+   token = stbi__hdr_gettoken(s,buffer);
7571+   if (strncmp(token, "-Y ", 3)) {
7572+       stbi__rewind( s );
7573+       return 0;
7574+   }
7575+   token += 3;
7576+   *y = (int) strtol(token, &token, 10);
7577+   while (*token == ' ') ++token;
7578+   if (strncmp(token, "+X ", 3)) {
7579+       stbi__rewind( s );
7580+       return 0;
7581+   }
7582+   token += 3;
7583+   *x = (int) strtol(token, NULL, 10);
7584+   *comp = 3;
7585+   return 1;
7586+}
7587+#endif /* STBI_NO_HDR */
7588+
7589+#ifndef STBI_NO_BMP
7590+static int stbi__bmp_info(stbi__context *s, int *x, int *y, int *comp)
7591+{
7592+   void *p;
7593+   stbi__bmp_data info;
7594+
7595+   info.all_a = 255;
7596+   p = stbi__bmp_parse_header(s, &info);
7597+   if (p == NULL) {
7598+      stbi__rewind( s );
7599+      return 0;
7600+   }
7601+   if (x) *x = s->img_x;
7602+   if (y) *y = s->img_y;
7603+   if (comp) {
7604+      if (info.bpp == 24 && info.ma == 0xff000000)
7605+         *comp = 3;
7606+      else
7607+         *comp = info.ma ? 4 : 3;
7608+   }
7609+   return 1;
7610+}
7611+#endif
7612+
7613+#ifndef STBI_NO_PSD
7614+static int stbi__psd_info(stbi__context *s, int *x, int *y, int *comp)
7615+{
7616+   int channelCount, dummy, depth;
7617+   if (!x) x = &dummy;
7618+   if (!y) y = &dummy;
7619+   if (!comp) comp = &dummy;
7620+   if (stbi__get32be(s) != 0x38425053) {
7621+       stbi__rewind( s );
7622+       return 0;
7623+   }
7624+   if (stbi__get16be(s) != 1) {
7625+       stbi__rewind( s );
7626+       return 0;
7627+   }
7628+   stbi__skip(s, 6);
7629+   channelCount = stbi__get16be(s);
7630+   if (channelCount < 0 || channelCount > 16) {
7631+       stbi__rewind( s );
7632+       return 0;
7633+   }
7634+   *y = stbi__get32be(s);
7635+   *x = stbi__get32be(s);
7636+   depth = stbi__get16be(s);
7637+   if (depth != 8 && depth != 16) {
7638+       stbi__rewind( s );
7639+       return 0;
7640+   }
7641+   if (stbi__get16be(s) != 3) {
7642+       stbi__rewind( s );
7643+       return 0;
7644+   }
7645+   *comp = 4;
7646+   return 1;
7647+}
7648+
7649+static int stbi__psd_is16(stbi__context *s)
7650+{
7651+   int channelCount, depth;
7652+   if (stbi__get32be(s) != 0x38425053) {
7653+       stbi__rewind( s );
7654+       return 0;
7655+   }
7656+   if (stbi__get16be(s) != 1) {
7657+       stbi__rewind( s );
7658+       return 0;
7659+   }
7660+   stbi__skip(s, 6);
7661+   channelCount = stbi__get16be(s);
7662+   if (channelCount < 0 || channelCount > 16) {
7663+       stbi__rewind( s );
7664+       return 0;
7665+   }
7666+   STBI_NOTUSED(stbi__get32be(s));
7667+   STBI_NOTUSED(stbi__get32be(s));
7668+   depth = stbi__get16be(s);
7669+   if (depth != 16) {
7670+       stbi__rewind( s );
7671+       return 0;
7672+   }
7673+   return 1;
7674+}
7675+#endif
7676+
7677+#ifndef STBI_NO_PIC
7678+static int stbi__pic_info(stbi__context *s, int *x, int *y, int *comp)
7679+{
7680+   int act_comp=0,num_packets=0,chained,dummy;
7681+   stbi__pic_packet packets[10];
7682+
7683+   if (!x) x = &dummy;
7684+   if (!y) y = &dummy;
7685+   if (!comp) comp = &dummy;
7686+
7687+   if (!stbi__pic_is4(s,"\x53\x80\xF6\x34")) {
7688+      stbi__rewind(s);
7689+      return 0;
7690+   }
7691+
7692+   stbi__skip(s, 88);
7693+
7694+   *x = stbi__get16be(s);
7695+   *y = stbi__get16be(s);
7696+   if (stbi__at_eof(s)) {
7697+      stbi__rewind( s);
7698+      return 0;
7699+   }
7700+   if ( (*x) != 0 && (1 << 28) / (*x) < (*y)) {
7701+      stbi__rewind( s );
7702+      return 0;
7703+   }
7704+
7705+   stbi__skip(s, 8);
7706+
7707+   do {
7708+      stbi__pic_packet *packet;
7709+
7710+      if (num_packets==sizeof(packets)/sizeof(packets[0]))
7711+         return 0;
7712+
7713+      packet = &packets[num_packets++];
7714+      chained = stbi__get8(s);
7715+      packet->size    = stbi__get8(s);
7716+      packet->type    = stbi__get8(s);
7717+      packet->channel = stbi__get8(s);
7718+      act_comp |= packet->channel;
7719+
7720+      if (stbi__at_eof(s)) {
7721+          stbi__rewind( s );
7722+          return 0;
7723+      }
7724+      if (packet->size != 8) {
7725+          stbi__rewind( s );
7726+          return 0;
7727+      }
7728+   } while (chained);
7729+
7730+   *comp = (act_comp & 0x10 ? 4 : 3);
7731+
7732+   return 1;
7733+}
7734+#endif
7735+
7736+/***************************************************************************************************/
7737+/*
7738+ * Portable Gray Map and Portable Pixel Map loader
7739+ * by Ken Miller
7740+ *
7741+ * PGM: http://netpbm.sourceforge.net/doc/pgm.html
7742+ * PPM: http://netpbm.sourceforge.net/doc/ppm.html
7743+ *
7744+ * Known limitations:
7745+ *    Does not support comments in the header section
7746+ *    Does not support ASCII image data (formats P2 and P3)
7747+ */
7748+
7749+#ifndef STBI_NO_PNM
7750+
7751+static int      stbi__pnm_test(stbi__context *s)
7752+{
7753+   char p, t;
7754+   p = (char) stbi__get8(s);
7755+   t = (char) stbi__get8(s);
7756+   if (p != 'P' || (t != '5' && t != '6')) {
7757+       stbi__rewind( s );
7758+       return 0;
7759+   }
7760+   return 1;
7761+}
7762+
7763+static void *stbi__pnm_load(stbi__context *s, int *x, int *y, int *comp, int req_comp, stbi__result_info *ri)
7764+{
7765+   stbi_uc *out;
7766+   STBI_NOTUSED(ri);
7767+
7768+   ri->bits_per_channel = stbi__pnm_info(s, (int *)&s->img_x, (int *)&s->img_y, (int *)&s->img_n);
7769+   if (ri->bits_per_channel == 0)
7770+      return 0;
7771+
7772+   if (s->img_y > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
7773+   if (s->img_x > STBI_MAX_DIMENSIONS) return stbi__errpuc("too large","Very large image (corrupt?)");
7774+
7775+   *x = s->img_x;
7776+   *y = s->img_y;
7777+   if (comp) *comp = s->img_n;
7778+
7779+   if (!stbi__mad4sizes_valid(s->img_n, s->img_x, s->img_y, ri->bits_per_channel / 8, 0))
7780+      return stbi__errpuc("too large", "PNM too large");
7781+
7782+   out = (stbi_uc *) stbi__malloc_mad4(s->img_n, s->img_x, s->img_y, ri->bits_per_channel / 8, 0);
7783+   if (!out) return stbi__errpuc("outofmem", "Out of memory");
7784+   if (!stbi__getn(s, out, s->img_n * s->img_x * s->img_y * (ri->bits_per_channel / 8))) {
7785+      STBI_FREE(out);
7786+      return stbi__errpuc("bad PNM", "PNM file truncated");
7787+   }
7788+
7789+   if (req_comp && req_comp != s->img_n) {
7790+      if (ri->bits_per_channel == 16) {
7791+         out = (stbi_uc *) stbi__convert_format16((stbi__uint16 *) out, s->img_n, req_comp, s->img_x, s->img_y);
7792+      } else {
7793+         out = stbi__convert_format(out, s->img_n, req_comp, s->img_x, s->img_y);
7794+      }
7795+      if (out == NULL) return out; /* stbi__convert_format frees input on failure */
7796+   }
7797+   return out;
7798+}
7799+
7800+static int      stbi__pnm_isspace(char c)
7801+{
7802+   return c == ' ' || c == '\t' || c == '\n' || c == '\v' || c == '\f' || c == '\r';
7803+}
7804+
7805+static void     stbi__pnm_skip_whitespace(stbi__context *s, char *c)
7806+{
7807+   for (;;) {
7808+      while (!stbi__at_eof(s) && stbi__pnm_isspace(*c))
7809+         *c = (char) stbi__get8(s);
7810+
7811+      if (stbi__at_eof(s) || *c != '#')
7812+         break;
7813+
7814+      while (!stbi__at_eof(s) && *c != '\n' && *c != '\r' )
7815+         *c = (char) stbi__get8(s);
7816+   }
7817+}
7818+
7819+static int      stbi__pnm_isdigit(char c)
7820+{
7821+   return c >= '0' && c <= '9';
7822+}
7823+
7824+static int      stbi__pnm_getinteger(stbi__context *s, char *c)
7825+{
7826+   int value = 0;
7827+
7828+   while (!stbi__at_eof(s) && stbi__pnm_isdigit(*c)) {
7829+      value = value*10 + (*c - '0');
7830+      *c = (char) stbi__get8(s);
7831+      if((value > 214748364) || (value == 214748364 && *c > '7'))
7832+          return stbi__err("integer parse overflow", "Parsing an integer in the PPM header overflowed a 32-bit int");
7833+   }
7834+
7835+   return value;
7836+}
7837+
7838+static int      stbi__pnm_info(stbi__context *s, int *x, int *y, int *comp)
7839+{
7840+   int maxv, dummy;
7841+   char c, p, t;
7842+
7843+   if (!x) x = &dummy;
7844+   if (!y) y = &dummy;
7845+   if (!comp) comp = &dummy;
7846+
7847+   stbi__rewind(s);
7848+
7849+   /* Get identifier */
7850+   p = (char) stbi__get8(s);
7851+   t = (char) stbi__get8(s);
7852+   if (p != 'P' || (t != '5' && t != '6')) {
7853+       stbi__rewind(s);
7854+       return 0;
7855+   }
7856+
7857+   *comp = (t == '6') ? 3 : 1; /* '5' is 1-component .pgm; '6' is 3-component .ppm */
7858+
7859+   c = (char) stbi__get8(s);
7860+   stbi__pnm_skip_whitespace(s, &c);
7861+
7862+   *x = stbi__pnm_getinteger(s, &c); /* read width */
7863+   if(*x == 0)
7864+       return stbi__err("invalid width", "PPM image header had zero or overflowing width");
7865+   stbi__pnm_skip_whitespace(s, &c);
7866+
7867+   *y = stbi__pnm_getinteger(s, &c); /* read height */
7868+   if (*y == 0)
7869+       return stbi__err("invalid width", "PPM image header had zero or overflowing width");
7870+   stbi__pnm_skip_whitespace(s, &c);
7871+
7872+   maxv = stbi__pnm_getinteger(s, &c); /* read max value */
7873+   if (maxv > 65535)
7874+      return stbi__err("max value > 65535", "PPM image supports only 8-bit and 16-bit images");
7875+   else if (maxv > 255)
7876+      return 16;
7877+   else
7878+      return 8;
7879+}
7880+
7881+static int stbi__pnm_is16(stbi__context *s)
7882+{
7883+   if (stbi__pnm_info(s, NULL, NULL, NULL) == 16)
7884+	   return 1;
7885+   return 0;
7886+}
7887+#endif
7888+
7889+static int stbi__info_main(stbi__context *s, int *x, int *y, int *comp)
7890+{
7891+   #ifndef STBI_NO_JPEG
7892+   if (stbi__jpeg_info(s, x, y, comp)) return 1;
7893+   #endif
7894+
7895+   #ifndef STBI_NO_PNG
7896+   if (stbi__png_info(s, x, y, comp))  return 1;
7897+   #endif
7898+
7899+   #ifndef STBI_NO_GIF
7900+   if (stbi__gif_info(s, x, y, comp))  return 1;
7901+   #endif
7902+
7903+   #ifndef STBI_NO_BMP
7904+   if (stbi__bmp_info(s, x, y, comp))  return 1;
7905+   #endif
7906+
7907+   #ifndef STBI_NO_PSD
7908+   if (stbi__psd_info(s, x, y, comp))  return 1;
7909+   #endif
7910+
7911+   #ifndef STBI_NO_PIC
7912+   if (stbi__pic_info(s, x, y, comp))  return 1;
7913+   #endif
7914+
7915+   #ifndef STBI_NO_PNM
7916+   if (stbi__pnm_info(s, x, y, comp))  return 1;
7917+   #endif
7918+
7919+   #ifndef STBI_NO_HDR
7920+   if (stbi__hdr_info(s, x, y, comp))  return 1;
7921+   #endif
7922+
7923+   /* test tga last because it's a crappy test! */
7924+   #ifndef STBI_NO_TGA
7925+   if (stbi__tga_info(s, x, y, comp))
7926+       return 1;
7927+   #endif
7928+   return stbi__err("unknown image type", "Image not of any known type, or corrupt");
7929+}
7930+
7931+static int stbi__is_16_main(stbi__context *s)
7932+{
7933+   #ifndef STBI_NO_PNG
7934+   if (stbi__png_is16(s))  return 1;
7935+   #endif
7936+
7937+   #ifndef STBI_NO_PSD
7938+   if (stbi__psd_is16(s))  return 1;
7939+   #endif
7940+
7941+   #ifndef STBI_NO_PNM
7942+   if (stbi__pnm_is16(s))  return 1;
7943+   #endif
7944+   return 0;
7945+}
7946+
7947+#ifndef STBI_NO_STDIO
7948+STBIDEF int stbi_info(char const *filename, int *x, int *y, int *comp)
7949+{
7950+    FILE *f = stbi__fopen(filename, "rb");
7951+    int result;
7952+    if (!f) return stbi__err("can't fopen", "Unable to open file");
7953+    result = stbi_info_from_file(f, x, y, comp);
7954+    fclose(f);
7955+    return result;
7956+}
7957+
7958+STBIDEF int stbi_info_from_file(FILE *f, int *x, int *y, int *comp)
7959+{
7960+   int r;
7961+   stbi__context s;
7962+   long pos = ftell(f);
7963+   stbi__start_file(&s, f);
7964+   r = stbi__info_main(&s,x,y,comp);
7965+   fseek(f,pos,SEEK_SET);
7966+   return r;
7967+}
7968+
7969+STBIDEF int stbi_is_16_bit(char const *filename)
7970+{
7971+    FILE *f = stbi__fopen(filename, "rb");
7972+    int result;
7973+    if (!f) return stbi__err("can't fopen", "Unable to open file");
7974+    result = stbi_is_16_bit_from_file(f);
7975+    fclose(f);
7976+    return result;
7977+}
7978+
7979+STBIDEF int stbi_is_16_bit_from_file(FILE *f)
7980+{
7981+   int r;
7982+   stbi__context s;
7983+   long pos = ftell(f);
7984+   stbi__start_file(&s, f);
7985+   r = stbi__is_16_main(&s);
7986+   fseek(f,pos,SEEK_SET);
7987+   return r;
7988+}
7989+#endif /* !STBI_NO_STDIO */
7990+
7991+STBIDEF int stbi_info_from_memory(stbi_uc const *buffer, int len, int *x, int *y, int *comp)
7992+{
7993+   stbi__context s;
7994+   stbi__start_mem(&s,buffer,len);
7995+   return stbi__info_main(&s,x,y,comp);
7996+}
7997+
7998+STBIDEF int stbi_info_from_callbacks(stbi_io_callbacks const *c, void *user, int *x, int *y, int *comp)
7999+{
8000+   stbi__context s;
8001+   stbi__start_callbacks(&s, (stbi_io_callbacks *) c, user);
8002+   return stbi__info_main(&s,x,y,comp);
8003+}
8004+
8005+STBIDEF int stbi_is_16_bit_from_memory(stbi_uc const *buffer, int len)
8006+{
8007+   stbi__context s;
8008+   stbi__start_mem(&s,buffer,len);
8009+   return stbi__is_16_main(&s);
8010+}
8011+
8012+STBIDEF int stbi_is_16_bit_from_callbacks(stbi_io_callbacks const *c, void *user)
8013+{
8014+   stbi__context s;
8015+   stbi__start_callbacks(&s, (stbi_io_callbacks *) c, user);
8016+   return stbi__is_16_main(&s);
8017+}
8018+
8019+#endif /* STB_IMAGE_IMPLEMENTATION */
8020+
8021+/*
8022+   revision history:
8023+      2.20  (2019-02-07) support utf8 filenames in Windows; fix warnings and platform ifdefs
8024+      2.19  (2018-02-11) fix warning
8025+      2.18  (2018-01-30) fix warnings
8026+      2.17  (2018-01-29) change sbti__shiftsigned to avoid clang -O2 bug
8027+                         1-bit BMP
8028+                         *_is_16_bit api
8029+                         avoid warnings
8030+      2.16  (2017-07-23) all functions have 16-bit variants;
8031+                         STBI_NO_STDIO works again;
8032+                         compilation fixes;
8033+                         fix rounding in unpremultiply;
8034+                         optimize vertical flip;
8035+                         disable raw_len validation;
8036+                         documentation fixes
8037+      2.15  (2017-03-18) fix png-1,2,4 bug; now all Imagenet JPGs decode;
8038+                         warning fixes; disable run-time SSE detection on gcc;
8039+                         uniform handling of optional "return" values;
8040+                         thread-safe initialization of zlib tables
8041+      2.14  (2017-03-03) remove deprecated STBI_JPEG_OLD; fixes for Imagenet JPGs
8042+      2.13  (2016-11-29) add 16-bit API, only supported for PNG right now
8043+      2.12  (2016-04-02) fix typo in 2.11 PSD fix that caused crashes
8044+      2.11  (2016-04-02) allocate large structures on the stack
8045+                         remove white matting for transparent PSD
8046+                         fix reported channel count for PNG & BMP
8047+                         re-enable SSE2 in non-gcc 64-bit
8048+                         support RGB-formatted JPEG
8049+                         read 16-bit PNGs (only as 8-bit)
8050+      2.10  (2016-01-22) avoid warning introduced in 2.09 by STBI_REALLOC_SIZED
8051+      2.09  (2016-01-16) allow comments in PNM files
8052+                         16-bit-per-pixel TGA (not bit-per-component)
8053+                         info() for TGA could break due to .hdr handling
8054+                         info() for BMP to shares code instead of sloppy parse
8055+                         can use STBI_REALLOC_SIZED if allocator doesn't support realloc
8056+                         code cleanup
8057+      2.08  (2015-09-13) fix to 2.07 cleanup, reading RGB PSD as RGBA
8058+      2.07  (2015-09-13) fix compiler warnings
8059+                         partial animated GIF support
8060+                         limited 16-bpc PSD support
8061+                         #ifdef unused functions
8062+                         bug with < 92 byte PIC,PNM,HDR,TGA
8063+      2.06  (2015-04-19) fix bug where PSD returns wrong '*comp' value
8064+      2.05  (2015-04-19) fix bug in progressive JPEG handling, fix warning
8065+      2.04  (2015-04-15) try to re-enable SIMD on MinGW 64-bit
8066+      2.03  (2015-04-12) extra corruption checking (mmozeiko)
8067+                         stbi_set_flip_vertically_on_load (nguillemot)
8068+                         fix NEON support; fix mingw support
8069+      2.02  (2015-01-19) fix incorrect assert, fix warning
8070+      2.01  (2015-01-17) fix various warnings; suppress SIMD on gcc 32-bit without -msse2
8071+      2.00b (2014-12-25) fix STBI_MALLOC in progressive JPEG
8072+      2.00  (2014-12-25) optimize JPG, including x86 SSE2 & NEON SIMD (ryg)
8073+                         progressive JPEG (stb)
8074+                         PGM/PPM support (Ken Miller)
8075+                         STBI_MALLOC,STBI_REALLOC,STBI_FREE
8076+                         GIF bugfix -- seemingly never worked
8077+                         STBI_NO_*, STBI_ONLY_*
8078+      1.48  (2014-12-14) fix incorrectly-named assert()
8079+      1.47  (2014-12-14) 1/2/4-bit PNG support, both direct and paletted (Omar Cornut & stb)
8080+                         optimize PNG (ryg)
8081+                         fix bug in interlaced PNG with user-specified channel count (stb)
8082+      1.46  (2014-08-26)
8083+              fix broken tRNS chunk (colorkey-style transparency) in non-paletted PNG
8084+      1.45  (2014-08-16)
8085+              fix MSVC-ARM internal compiler error by wrapping malloc
8086+      1.44  (2014-08-07)
8087+              various warning fixes from Ronny Chevalier
8088+      1.43  (2014-07-15)
8089+              fix MSVC-only compiler problem in code changed in 1.42
8090+      1.42  (2014-07-09)
8091+              don't define _CRT_SECURE_NO_WARNINGS (affects user code)
8092+              fixes to stbi__cleanup_jpeg path
8093+              added STBI_ASSERT to avoid requiring assert.h
8094+      1.41  (2014-06-25)
8095+              fix search&replace from 1.36 that messed up comments/error messages
8096+      1.40  (2014-06-22)
8097+              fix gcc struct-initialization warning
8098+      1.39  (2014-06-15)
8099+              fix to TGA optimization when req_comp != number of components in TGA;
8100+              fix to GIF loading because BMP wasn't rewinding (whoops, no GIFs in my test suite)
8101+              add support for BMP version 5 (more ignored fields)
8102+      1.38  (2014-06-06)
8103+              suppress MSVC warnings on integer casts truncating values
8104+              fix accidental rename of 'skip' field of I/O
8105+      1.37  (2014-06-04)
8106+              remove duplicate typedef
8107+      1.36  (2014-06-03)
8108+              convert to header file single-file library
8109+              if de-iphone isn't set, load iphone images color-swapped instead of returning NULL
8110+      1.35  (2014-05-27)
8111+              various warnings
8112+              fix broken STBI_SIMD path
8113+              fix bug where stbi_load_from_file no longer left file pointer in correct place
8114+              fix broken non-easy path for 32-bit BMP (possibly never used)
8115+              TGA optimization by Arseny Kapoulkine
8116+      1.34  (unknown)
8117+              use STBI_NOTUSED in stbi__resample_row_generic(), fix one more leak in tga failure case
8118+      1.33  (2011-07-14)
8119+              make stbi_is_hdr work in STBI_NO_HDR (as specified), minor compiler-friendly improvements
8120+      1.32  (2011-07-13)
8121+              support for "info" function for all supported filetypes (SpartanJ)
8122+      1.31  (2011-06-20)
8123+              a few more leak fixes, bug in PNG handling (SpartanJ)
8124+      1.30  (2011-06-11)
8125+              added ability to load files via callbacks to accomidate custom input streams (Ben Wenger)
8126+              removed deprecated format-specific test/load functions
8127+              removed support for installable file formats (stbi_loader) -- would have been broken for IO callbacks anyway
8128+              error cases in bmp and tga give messages and don't leak (Raymond Barbiero, grisha)
8129+              fix inefficiency in decoding 32-bit BMP (David Woo)
8130+      1.29  (2010-08-16)
8131+              various warning fixes from Aurelien Pocheville
8132+      1.28  (2010-08-01)
8133+              fix bug in GIF palette transparency (SpartanJ)
8134+      1.27  (2010-08-01)
8135+              cast-to-stbi_uc to fix warnings
8136+      1.26  (2010-07-24)
8137+              fix bug in file buffering for PNG reported by SpartanJ
8138+      1.25  (2010-07-17)
8139+              refix trans_data warning (Won Chun)
8140+      1.24  (2010-07-12)
8141+              perf improvements reading from files on platforms with lock-heavy fgetc()
8142+              minor perf improvements for jpeg
8143+              deprecated type-specific functions so we'll get feedback if they're needed
8144+              attempt to fix trans_data warning (Won Chun)
8145+      1.23    fixed bug in iPhone support
8146+      1.22  (2010-07-10)
8147+              removed image *writing* support
8148+              stbi_info support from Jetro Lauha
8149+              GIF support from Jean-Marc Lienher
8150+              iPhone PNG-extensions from James Brown
8151+              warning-fixes from Nicolas Schulz and Janez Zemva (i.stbi__err. Janez (U+017D)emva)
8152+      1.21    fix use of 'stbi_uc' in header (reported by jon blow)
8153+      1.20    added support for Softimage PIC, by Tom Seddon
8154+      1.19    bug in interlaced PNG corruption check (found by ryg)
8155+      1.18  (2008-08-02)
8156+              fix a threading bug (local mutable static)
8157+      1.17    support interlaced PNG
8158+      1.16    major bugfix - stbi__convert_format converted one too many pixels
8159+      1.15    initialize some fields for thread safety
8160+      1.14    fix threadsafe conversion bug
8161+              header-file-only version (#define STBI_HEADER_FILE_ONLY before including)
8162+      1.13    threadsafe
8163+      1.12    const qualifiers in the API
8164+      1.11    Support installable IDCT, colorspace conversion routines
8165+      1.10    Fixes for 64-bit (don't use "unsigned long")
8166+              optimized upsampling by Fabian "ryg" Giesen
8167+      1.09    Fix format-conversion for PSD code (bad global variables!)
8168+      1.08    Thatcher Ulrich's PSD code integrated by Nicolas Schulz
8169+      1.07    attempt to fix C++ warning/errors again
8170+      1.06    attempt to fix C++ warning/errors again
8171+      1.05    fix TGA loading to return correct *comp and use good luminance calc
8172+      1.04    default float alpha is 1, not 255; use 'void *' for stbi_image_free
8173+      1.03    bugfixes to STBI_NO_STDIO, STBI_NO_HDR
8174+      1.02    support for (subset of) HDR files, float interface for preferred access to them
8175+      1.01    fix bug: possible bug in handling right-side up bmps... not sure
8176+              fix bug: the stbi__bmp_load() and stbi__tga_load() functions didn't work at all
8177+      1.00    interface to zlib that skips zlib header
8178+      0.99    correct handling of alpha in palette
8179+      0.98    TGA loader by lonesock; dynamically add loaders (untested)
8180+      0.97    jpeg errors on too large a file; also catch another malloc failure
8181+      0.96    fix detection of invalid v value - particleman@mollyrocket forum
8182+      0.95    during header scan, seek to markers in case of padding
8183+      0.94    STBI_NO_STDIO to disable stdio usage; rename all #defines the same
8184+      0.93    handle jpegtran output; verbose errors
8185+      0.92    read 4,8,16,24,32-bit BMP files of several formats
8186+      0.91    output 24-bit Windows 3.0 BMP files
8187+      0.90    fix a few more warnings; bump version number to approach 1.0
8188+      0.61    bugfixes due to Marc LeBlanc, Christopher Lloyd
8189+      0.60    fix compiling as c++
8190+      0.59    fix warnings: merge Dave Moore's -Wall fixes
8191+      0.58    fix bug: zlib uncompressed mode len/nlen was wrong endian
8192+      0.57    fix bug: jpg last huffman symbol before marker was >9 bits but less than 16 available
8193+      0.56    fix bug: zlib uncompressed mode len vs. nlen
8194+      0.55    fix bug: restart_interval not initialized to 0
8195+      0.54    allow NULL for 'int *comp'
8196+      0.53    fix bug in png 3->4; speedup png decoding
8197+      0.52    png handles req_comp=3,4 directly; minor cleanup; jpeg comments
8198+      0.51    obey req_comp requests, 1-component jpegs return as 1-component,
8199+              on 'test' only check type, not whether we support this variant
8200+      0.50  (2006-11-19)
8201+              first released version
8202+*/
8203+
8204+
8205+/*
8206+------------------------------------------------------------------------------
8207+This software is available under 2 licenses -- choose whichever you prefer.
8208+------------------------------------------------------------------------------
8209+ALTERNATIVE A - MIT License
8210+Copyright (c) 2017 Sean Barrett
8211+Permission is hereby granted, free of charge, to any person obtaining a copy of
8212+this software and associated documentation files (the "Software"), to deal in
8213+the Software without restriction, including without limitation the rights to
8214+use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies
8215+of the Software, and to permit persons to whom the Software is furnished to do
8216+so, subject to the following conditions:
8217+The above copyright notice and this permission notice shall be included in all
8218+copies or substantial portions of the Software.
8219+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
8220+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
8221+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
8222+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
8223+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
8224+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
8225+SOFTWARE.
8226+------------------------------------------------------------------------------
8227+ALTERNATIVE B - Public Domain (www.unlicense.org)
8228+This is free and unencumbered software released into the public domain.
8229+Anyone is free to copy, modify, publish, use, compile, sell, or distribute this
8230+software, either in source code form or as a compiled binary, for any purpose,
8231+commercial or non-commercial, and by any means.
8232+In jurisdictions that recognize copyright laws, the author or authors of this
8233+software dedicate any and all copyright interest in the software to the public
8234+domain. We make this dedication for the benefit of the public at large and to
8235+the detriment of our heirs and successors. We intend this dedication to be an
8236+overt act of relinquishment in perpetuity of all present and future rights to
8237+this software under copyright law.
8238+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
8239+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
8240+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
8241+AUTHORS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN
8242+ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
8243+WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
8244+------------------------------------------------------------------------------
8245+*/