Machine-readable medium and data

ISBN, a unique numeric book identifier, represented as an EAN-13 bar code. Showing both machine-readable bars, and human-readable digits.

In communications and computing, a machine-readable medium (or computer-readable medium) is a medium capable of storing data in a format easily readable by a digital computer or a sensor. It contrasts with human-readable medium and data.

The result is called machine-readable data or computer-readable data, and the data itself can be described as having machine-readability.

Data

Machine-readable data must be structured data.[1]

Attempts to create machine-readable data occurred as early as the 1960s. At the same time that seminal developments in machine-reading and natural-language processing were releasing (like Weizenbaum'sELIZA), people were anticipating the success of machine-readable functionality and attempting to create machine-readable documents. One such example was musicologist Nancy B. Reich's creation of a machine-readable catalog of composer William Jay Sydeman's works in 1966.

In the United States, the OPEN Government Data Act of 14 January 2019 defines machine-readable data as "data in a format that can be easily processed by a computer without human intervention while ensuring no semantic meaning is lost." The law directs U.S. federal agencies to publish public data in such a manner,[2] ensuring that "any public data asset of the agency is machine-readable".[3]

Machine-readable data may be classified into two groups: human-readable data that is marked up so that it can also be read by machines (e.g. microformats, RDFa, HTML), and data file formats intended principally for processing by machines (CSV, RDF, XML, JSON). These formats are only machine readable if the data contained within them is formally structured; exporting a CSV file from a badly structured spreadsheet does not meet the definition.

Machine readable is not synonymous with digitally accessible. A digitally accessible document may be online, making it easier for humans to access via computers, but its content is much harder to extract, transform, and process via computer programming logic if it is not machine-readable.[4]

Extensible Markup Language (XML) is designed to be both human- and machine-readable, and Extensible Stylesheet Language Transformations (XSLT) is used to improve the presentation of the data for human readability. For example, XSLT can be used to automatically render XML in Portable Document Format (PDF). Machine-readable data can be automatically transformed for human-readability but, generally speaking, the reverse is not true.

For purposes of implementation of the Government Performance and Results Act (GPRA) Modernization Act, the Office of Management and Budget (OMB) defines "machine readable format" as follows: "Format in a standard computer language (not English text) that can be read automatically by a web browser or computer system. (e.g.; xml). Traditional word processing documents and portable document format (PDF) files are easily read by humans but typically are difficult for machines to interpret. Other formats such as extensible markup language (XML), (JSON), or spreadsheets with header columns that can be exported as comma separated values (CSV) are machine readable formats. As HTML is a structural markup language, discreetly labeling parts of the document, computers are able to gather document components to assemble tables of contents, outlines, literature search bibliographies, etc. It is possible to make traditional word processing documents and other formats machine readable but the documents must include enhanced structural elements."[5]

Media

Examples of machine-readable media include magnetic media such as magnetic disks, cards, tapes, and drums, punched cards and paper tapes, optical discs, barcodes and magnetic ink characters.

Common machine-readable technologies include magnetic recording, processing waveforms, and barcodes. Optical character recognition (OCR) can be used to enable machines to read information available to humans. Any information retrievable by any form of energy can be machine-readable.

Examples include:

Applications

Documents

A machine-readable document is a document whose content can be readily processed by computers. Such documents are distinguished from more general machine-readable data by virtue of having further structure to provide the necessary context to support the business processes for which they are created.

Catalogs

MARC (machine-readable cataloging) is a standard set of digitalformats for the machine-readable description of items catalogued by libraries, such as books, DVDs, and digital resources. Computerized library catalogs and library management software need to structure their catalog records as per an industry-wide standard, which is MARC, so that bibliographic information can be shared freely between computers. The structure of bibliographic records almost universally follows the MARC standard. Other standards work in conjunction with MARC, for example, Anglo-American Cataloguing Rules (AACR)/Resource Description and Access (RDA) provide guidelines on formulating bibliographic data into the MARC record structure, while the International Standard Bibliographic Description (ISBD) provides guidelines for displaying MARC records in a standard, human-readable form.

Dictionaries

Machine-readable dictionary (MRD) is a dictionary stored as machine-readable data instead of being printed on paper. It is an electronic dictionary and lexical database.

القاموس المقروء آليًا هو قاموس إلكتروني يُمكن تحميله في قاعدة بيانات والاستعلام عنه عبر برامج تطبيقية. قد يكون قاموسًا تفسيريًا أحادي اللغة، أو قاموسًا متعدد اللغات لدعم الترجمة بين لغتين أو أكثر، أو مزيجًا من الاثنين. عادةً ما تستخدم برامج الترجمة بين لغات متعددة قواميس ثنائية الاتجاه. قد يكون القاموس المقروء آليًا قاموسًا ذا بنية خاصة يُستعلم عنه بواسطة برامج مخصصة (عبر الإنترنت مثلاً)، أو قد يكون قاموسًا ذا بنية مفتوحة متاحًا للتحميل في قواعد بيانات الحاسوب، وبالتالي يُمكن استخدامه عبر تطبيقات برمجية متنوعة. تحتوي القواميس التقليدية على جذر مع أوصاف مختلفة. قد يتمتع القاموس المقروء آليًا بقدرات إضافية، ولذلك يُطلق عليه أحيانًا اسم القاموس الذكي. مثال على القاموس الذكي هو قاموس Gellish الإنجليزي مفتوح المصدر .

يُستخدم مصطلح "قاموس" أيضًا للإشارة إلى معجم إلكتروني ، كما هو الحال في برامج التدقيق الإملائي . إذا رُتِّبت القواميس وفقًا لتسلسل هرمي للمفاهيم (أو المصطلحات) من نوع فرعي إلى نوع رئيسي، يُطلق عليها اسم " تصنيف" . أما إذا احتوت على علاقات أخرى بين المفاهيم، فتُسمى " أنطولوجيا" . قد تستخدم محركات البحث إما المعجم أو التصنيف أو الأنطولوجيا لتحسين نتائج البحث. ومن أنواع القواميس الإلكترونية المتخصصة: القواميس الصرفية والقواميس النحوية.

يُستخدم مصطلح MRD غالبًا في مقابل قاموس NLP ، حيث يُمثل MRD النسخة الإلكترونية من القاموس الذي كان مطبوعًا ورقيًا. ورغم استخدام كلا المصطلحين في البرامج، إلا أن مصطلح قاموس NLP هو المُفضل عندما يُبنى القاموس من الصفر مع مراعاة متطلبات معالجة اللغة الطبيعية. ويُمكن لمعيار ISO الخاص بـ MRD وNLP تمثيل كلا البنيتين، ويُسمى إطار عمل الترميز المعجمي (Lexical Markup Framework ). [ 6 ]

جوازات السفر

جواز السفر المقروء آلياً (MRP) هو وثيقة سفر مقروءة آلياً (MRTD) تحتوي على بيانات صفحة الهوية مُشفّرة بتقنية التعرف الضوئي على الأحرف ( OCR). بدأت العديد من الدول بإصدار وثائق سفر مقروءة آلياً في ثمانينيات القرن الماضي. معظم جوازات السفر في العالم هي من هذا النوع. وقد ألزمت منظمة الطيران المدني الدولي (ICAO) جميع الدول الأعضاء فيها بإصدار جوازات سفر مقروءة آلياً فقط اعتباراً من 1 أبريل 2010، على أن تنتهي صلاحية جميع جوازات السفر غير المقروءة آلياً بحلول 24 نوفمبر 2015. [ 7 ]

تُعتمد جوازات السفر المقروءة آليًا وفقًا لوثيقة منظمة الطيران المدني الدولي رقم 9303 (المعتمدة من قبل المنظمة الدولية للتوحيد القياسي واللجنة الكهروتقنية الدولية تحت مسمى ISO/IEC 7501-1)، وتحتوي على منطقة خاصة للقراءة الآلية ( MRZ )، والتي عادةً ما تكون في أسفل صفحة البيانات الشخصية في بداية جواز السفر. وتصف وثيقة منظمة الطيران المدني الدولي رقم 9303 ثلاثة أنواع من الوثائق تتوافق مع أحجام ISO/IEC 7810 .

  • يُعدّ "النوع 3" نموذجياً لكتيبات جوازات السفر. ويتكون رمز القراءة الآلية (MRZ) من سطرين × 44 حرفاً.
  • "النوع 2" نادر نسبياً ويتكون من سطرين × 36 حرفاً.
  • "النوع 1" بحجم بطاقة الائتمان ويحتوي على 3 أسطر × 30 حرفًا.

يُتيح النموذج الثابت تحديد نوع الوثيقة، والاسم، ورقم الوثيقة، والجنسية، وتاريخ الميلاد، والجنس، وتاريخ انتهاء صلاحية الوثيقة. جميع هذه الحقول مطلوبة في جواز السفر. كما يُتاح مجال لإضافة معلومات إضافية اختيارية، تختلف باختلاف البلد. وتتوفر أيضًا نوعان من التأشيرات المقروءة آليًا بمواصفات مماثلة.

تستطيع أجهزة الكمبيوتر المزودة بكاميرا وبرامج مناسبة قراءة المعلومات الموجودة على جوازات السفر المقروءة آلياً بشكل مباشر. وهذا يُتيح معالجة أسرع للمسافرين القادمين من قِبل مسؤولي الهجرة، ودقة أكبر من جوازات السفر المقروءة يدوياً، فضلاً عن إدخال البيانات بشكل أسرع، وقراءة بيانات أكثر، ومطابقة أفضل للبيانات مع قواعد بيانات الهجرة وقوائم المراقبة.

إلى جانب المعلومات المقروءة ضوئيًا، تحتوي العديد من جوازات السفر على شريحة RFID تُمكّن أجهزة الكمبيوتر من قراءة كمية أكبر من المعلومات، مثل صورة حامل الجواز. تُسمى هذه الجوازات بجوازات السفر البيومترية ، وهي مُدرجة أيضًا في معيار منظمة الطيران المدني الدولي ICAO 9303.

انظر أيضاً

مراجع

  1. http://-07-22 .{{cite web}}: تحقق من |url=القيمة ( مساعدة ) ؛ مفقودة أو فارغة |title=( مساعدة )
  2. "HR4174" . stratml.us .
  3. "HR4174" . stratml.us .
  4. هيندلر، جيم؛ باردو، تيريزا أ. (24-09-2012). "مقدمة في قابلية قراءة المستندات والبيانات الإلكترونية آليًا" . Data.gov . مؤرشف من الأصل في 20-03-2021 . تم الاطلاع عليه في 27-02-2015 .
  5. تعميم مكتب الإدارة والميزانية رقم أ-11، الجزء 6 ، إعداد الميزانية وتقديمها وتنفيذها
  6. جيل فرانكوبولو (محرر) إطار عمل الترميز المعجمي LMF، ISTE / Wiley 2013 ( ISBN 978-1-84821-430-9)
  7. "الأسبوع الأخير للدول لضمان انتهاء صلاحية جوازات السفر غير المقروءة آلياً" . منظمة الطيران المدني الدولي (إيكاو) . مونتريال. 17 نوفمبر 2015. تاريخ الاطلاع: 11 مارس 2024 .

Public Domain تتضمن هذه المقالة موادًا متاحة للعموم من المعيار الفيدرالي 1037C ، إدارة الخدمات العامة . مؤرشفة من الأصل بتاريخ 22 يناير 2022.