Reading: Difference between revisions

Ahayri (talk | contribs)
Gorilli09 (talk | contribs)
Redirected page to Literature
Tag: New redirect
 
(4 intermediate revisions by one other user not shown)
Line 1: Line 1:
{{Stub}}
#REDIRECT[[Literature]]
Just like using [[Home_Media_Player#Media_player_software|media player software]] to access your home media library on a PC instead of other physical devices, your computer can be a versatile platform for reading a wide range of digital documents. This includes books and novels, magazines, catalogs, pamphlets and brochures, posters, reports, presentations, booklets, manuals, restaurant menus, newsletters, comics, and manga. This however, shouldn't be confused with ebooks, see [[E-reader]] page for that. Similar to the [[Ripping_games|disc ripping]] process, scanning (properly) physical documents/papers with their artwork and covers is often less preferred due to the need for additional hardware, larger storage requirements, and other conveniences of digitization (OCR process and existence of ebooks) at the expense of proper [[Preservation_projects|preservation]].
 
==Proper preservation==
<blockquote>''"Best practice dictates that master files should be lossless. If there are storage space constraints, other members of the organization in addition to archivists or digitization staff may be consulted on the organizational policy for master file formats."'' <ref>https://www.digitizationguidelines.gov/guidelines/FADGI%20Technical%20Guidelines%20for%20Digitizing%20Cultural%20Heritage%20Materials_3rd%20Edition_05092023.pdf</ref></blockquote>
 
When it comes to proper scanning for the preservation of paper documents and archival materials, DPI (dots per inch) and image resolution alone are not sufficient indicators of quality or suitability for long-term preservation (just like transcoding and [[ripping_games|ripping, dumping]] processes). While these metrics are important, they are just one part of a broader set of considerations that include hardware quality, scan settings, encoding and compression options, and file formats. Proper preservation requires a holistic approach to ensure that digital surrogates are both faithful to the original documents and durable for future access.
 
;Hardware Quality
The quality of the scanning hardware is a critical factor. High-quality scanners, such as those designed for archival purposes, offer superior optics, precise color calibration, and consistent lighting to capture fine details accurately. For example, planetary scanners or flatbed scanners with CCD (charge-coupled device) sensors are often preferred for delicate or bound materials, as they minimize physical stress on originals. In contrast, consumer-grade scanners or sheet-fed scanners may introduce distortions, inconsistent lighting, or mechanical damage, particularly for fragile documents like manuscripts or photographs. Specialized equipment, such as overhead scanners or those with book cradles, is often necessary for rare or oversized items to avoid damage during the scanning process.
 
;Scan Settings
Beyond hardware, scan settings play a pivotal role in achieving preservation-quality digital images. Key settings include:
*Bit Depth: A higher bit depth (e.g., 16-bit per channel) captures a greater range of color and grayscale information, which is essential for preserving the nuances of historical documents or photographs. For text-only documents, 8-bit grayscale may suffice, but color documents or images benefit from higher bit depths.
*Resolution: While DPI is often overemphasized, it remains important. For archival purposes, a resolution of 300–600 DPI is typically recommended, depending on the document type. Fine details, such as those in maps or illustrations, may require 600 DPI or higher, while standard text documents may be adequately captured at 300 DPI. Oversampling (scanning at excessively high DPI) can lead to unnecessarily large file sizes without proportional gains in quality.
*Color Space: Using a standardized color space, such as sRGB or Adobe RGB, ensures accurate color representation. Archival scans should avoid proprietary or device-dependent color spaces to ensure compatibility and consistency across systems.
*File Format for Capture: Scanning directly to a lossless format like TIFF is ideal for preservation, as it avoids data loss during capture. TIFF files retain all image data, making them suitable for creating master copies that can be used for further processing or derivative creation.
 
;Encoding and Compression Options
The choice of encoding and compression significantly impacts the quality and longevity of digital files. For preservation, lossless compression or uncompressed formats are preferred to maintain fidelity to the original scan. Common options include:
*TIFF: Widely used for archival master files, TIFF supports lossless compression (e.g., LZW, ZIP) and preserves all image data. It is a stable, non-proprietary format suitable for long-term storage. However, there’s also TIFF files with JPEG compression, which should be avoided. CCITT Group 4 could be used for bitonal images which is a lossless method of image compression.
*FLIF: The Free Lossless Image Format (FLIF) is an emerging lossless format that aims to outperform PNG and other lossless formats in terms of compression efficiency. FLIF supports progressive decoding and high bit depths, making it a potential candidate for archival use. However, its adoption in archival communities is minimal due to its relative newness, limited software support, and lack of established long-term stability compared to TIFF.
*PNG: The Portable Network Graphics (PNG) format is another lossless option commonly used for digital preservation. It supports high bit depths and transparency, making it suitable for scanned documents with simple color profiles or line art. While PNG files can be smaller than uncompressed TIFFs due to efficient lossless compression, they are less commonly used in archival settings due to limited metadata support compared to TIFF.
*JPEG 2000: This format offers both lossless and lossy compression options. Its wavelet-based compression can reduce file sizes while maintaining high quality, making it a popular choice for large-scale digitization projects. However, its adoption is sometimes limited by software compatibility and the need for specialized tools.
*JPEG XR: JPEG XR (Extended Range) supports both lossless and lossy compression and is designed to handle high dynamic range images with better compression efficiency than traditional JPEG. It offers advantages like support for higher bit depths and alpha channels, but its use in archival communities is limited due to less widespread software support and concerns about long-term format stability compared to TIFF or PNG. Lossy JPEG XR files, if repeatedly re-compressed, suffer from generational loss similar to JPEG.
*JPEG: JPEG is common for access copies due to its small file size, but its lossy compression discards data (how much depends on the compression settings of course), which can degrade quality especially over multiple saves or conversions. For example, users who repeatedly save JPEG files (e.g., editing and re-saving an already compressed image) introduce generational loss, where each save further degrades details, colors, and sharpness. This is comparable to re-encoding a compressed video or audio file, where quality erodes with each iteration. For preservation, JPEG should be avoided for master copies, though it may be suitable for low-quality derivatives intended for web access.
*DjVu: DjVu is a specialized format designed for scanned documents, particularly those combining text, line art, and images. It uses advanced compression techniques to separate text and background layers, allowing high compression ratios while maintaining readability. DjVu is especially effective for documents with large amounts of text, such as books or journals, as it can produce smaller file sizes than TIFF or JPEG 2000 without significant loss of quality for text. However, its use in preservation is less common due to limited software support and a smaller user base compared to others. DjVu is best suited for access copies or specific use cases where file size is a critical concern, but TIFF remains the gold standard for archival master files.
*WebP: WebP is a modern image format that supports both lossless and lossy compression. Its lossless mode provides efficient compression, often producing smaller file sizes than PNG or TIFF, while maintaining image quality. However, WebP is primarily designed for web use and has limited adoption in archival contexts due to concerns about long-term format support and metadata capabilities. Repeated lossy compression of WebP files can degrade quality, much like JPEG or lossy JPEG 2000, making it unsuitable for master copies.
 
Proper compression ensures that file sizes remain manageable without sacrificing critical details, but lossless compression or uncompressed formats are essential for archival masters to avoid cumulative degradation over time. Users must avoid the common error of compressing already compressed files, particularly with lossy formats like JPEG, WebP, or JPEG XR, as this practice introduces irreversible quality loss that undermines preservation goals.
 
;File Containers
File containers like CBZ, CBR, or PDF serve as wrappers for digital images and have minimal impact on preservation quality, as they primarily organize and package the underlying image files. However, their choice can affect accessibility and metadata management:
*PDF: A widely supported format that can embed metadata, making it useful for organizing documents and ensuring discoverability. PDF/A, a subset of PDF designed for archival purposes, ensures long-term accessibility by embedding fonts and [[Copy_protection|avoiding features]] like encryption that could hinder future access.
*CBZ/CBR: These are comic book archive formats (which is essentially ZIP or RAR archives) that package images (typically JPEGs and JPEG 2000 or rarely TIFF) into a single file. While useful for organizing sequential images, they are less common in formal archival settings due to limited metadata support. For preservation, the container should support robust metadata standards (e.g., Dublin Core or METS) to ensure that contextual information, such as document provenance, creation date, and scanning parameters, is preserved alongside the images.
*DjVu as a Container: DjVu files can also function as containers, bundling multiple pages or images into a single file, similar to PDF. DjVu’s layered compression makes it efficient for distributing scanned documents, especially for online access, but its preservation use is limited by the same software compatibility issues as its image format. Like CBZ/CBR, it is more suited for access copies than archival masters.
 
See [[E-reader]] page for ebook file formats such as EPUB, FictionBook/FB2 or MOBI.
 
;Additional Considerations for Preservation
*Metadata: Embedding detailed metadata within files or in accompanying records is crucial for long-term preservation. Metadata should include information about the original document (e.g., title, creator, date), the scanning process (e.g., scanner model, settings, date of digitization), and rights information to ensure future usability.
*Storage and Redundancy: Preservation requires secure, redundant storage systems to protect against data loss. This includes maintaining multiple copies of master files in geographically dispersed locations, ideally on stable media like enterprise-grade hard drives or cloud storage with strong checksum validation to detect corruption. See [[Copy protection#Physical reliability vs. digital storage]] section for more information about this.
*Migration and Format Stability: Digital preservation involves planning for format migration to prevent obsolescence. Formats like TIFF and PDF/A are preferred for their stability and widespread support, but ongoing monitoring of file format standards is necessary to ensure long-term accessibility.
*Access Copies vs. Master Files: Preservation workflows typically involve creating high-quality master files (e.g., uncompressed TIFFs or PNGs) for long-term storage and lower-quality derivatives (e.g., JPEG, WebP, JPEG XR, or DjVu) for access. This balances preservation needs with practical considerations like storage space and ease of access for researchers or the public.
*Avoiding User Errors in Compression: A critical preservation concern is avoiding user errors that degrade quality, such as repeatedly compressing already compressed files. For example, re-saving a JPEG, lossy WebP, or lossy JPEG XR file multiple times introduces generational loss, much like transcoding a video or an audio file repeatedly. Each compression or encoding cycle discards additional data, resulting in artifacts, loss of detail, and reduced fidelity to the original. Even with formats like DjVu, repeated lossy compression can degrade image quality, particularly for non-text elements. To prevent this, preservation workflows should maintain uncompressed or lossless master files (e.g., TIFF or PNG) and only create lossy derivatives (e.g., JPEG, WebP, or DjVu) for access purposes, ensuring that the master file remains untouched by subsequent compressions.
 
;Best Practices and Standards
Adhering to standards like those from the International Organization for Standardization (ISO) or the Library of Congress’s Federal Agencies Digital Guidelines Initiative (FADGI) ensures consistency and quality. For example, FADGI recommends specific technical parameters for scanning, such as a minimum of 4000 pixels across the long dimension for most documents and specific color management protocols. These standards help ensure that digital surrogates are both accurate representations of originals and viable for future use.
 
==Analyzing image files==
Several tools allow users to inspect and analyze image file properties, such as format, compression method, resolution, and metadata (e.g., EXIF, IPTC, XMP), similar to [[Home_Media_Player#Analyzing_video_and_audio_files|video and audio file tools]] or [[Copy_protection#Disc ripping and scan for copy protection|disc scanning tools]]. These tools range from dedicated software to built-in operating system utilities. These tools provide accessible ways to analyze image file properties, e.g., detailed insights into TIFF compression methods (e.g., LZW, JPEG, PackBits, or none) and other technical attributes.
 
*XnView MP: A versatile, free image viewer and organizer for Windows, macOS, and Linux, supporting over 500 image formats, including TIFF, JPEG, and PNG. Its "Properties" panel provides detailed technical information, such as compression method (e.g., LZW, JPEG, or none for TIFF files), color depth, resolution, and metadata. Users can access this by right-clicking an image and selecting "Properties" or using the Info panel in thumbnail view. XnView MP also supports batch processing and metadata editing.
*IrfanView: A lightweight, free image viewer for Windows that supports TIFF, JPEG, and other formats. It displays file properties, including compression type, bit depth, and metadata, via the "Image > Information" menu. IrfanView is efficient for quick analysis and supports plugins for extended functionality, such as advanced TIFF handling.
*GIMP: An open-source image editing program for Windows, macOS, and Linux that supports TIFF and other formats. While primarily an editor, GIMP can display basic file properties, including compression details, through the "File > Properties" menu. It’s ideal for users needing both analysis and editing capabilities.
 
Built-in OS Tools: Most operating systems include basic image viewers with property inspection features. For example:
*Windows Photos (Windows): Right-click an image and select "Properties" or use the "Details" tab to view format, compression, and metadata like EXIF. Limited for advanced TIFF compression details.
*Preview (macOS): Open an image and use "Tools > Show Inspector" to view format, compression, and metadata. Supports TIFF and GeoTIFF files natively.
*Eye of GNOME or Gwenview (Linux): These default viewers display basic properties like format and resolution, with limited compression details.
 
==Software types==
;Document viewer desktop software
*[https://www.sumatrapdfreader.org/ SumatraPDF]
*[https://www.cdisplayex.com/ CDisplayEx]
*[https://github.com/binarynonsense/comic-book-reader ACBR]
 
;Websites
There are several websites out there replicating the experience of reading via focusing things like page turning animation, visual elements, sounds, and other interactive elements for better immersion and enhancing the reading experience compared to looking a 2d monitor screen using document viewer software. Kinda similar to [[Virtual reality#VR game room simulations|game room simulations]], it's all about improving the reading experience.
*[https://heyzine.com/ heyzine.com]
*[https://publuu.com/ publuu.com]
*[https://dearflip.com/pdf-viewer/ dearflip.com]
 
;Game Room Simulations
There are [[Virtual reality#VR game room simulations|game room simulations]] which some of them supports the comics, books etc.
 
==See also==
*[[The Visual Novel Problem]]
*[[Home_Media_Player#Media_player_software|Media player software]]
*[[Preservation projects]]
*[[E-reader]]
*[[ripping_games|Ripping and dumping]]
*[[Broadcast and cable communication systems]]
 
==External links==
*[[Wikipedia:Reading]]
*[https://www.digitizationguidelines.gov/guidelines/File_format_compare.html digitizationguidelines: File Format Comparison Projects]
*[https://github.com/binarynonsense/comic-book-reader/issues/169 feature request for immersive elements on ACBR]
 
[[Category:Miscellaneous software categories]]
[[Category:Preservation]]
[[Category:Not really emulators]]
[[Category:FAQs]]