Emerging Chapter 2 Questions
PaRa Freshman
0:00 / 0:00
Emerging Chapter 2 Questions
682 просмотра · 9 месяцев назад
PaRa Freshman
4,8 тыс. подписчиков
682 просмотра · 9 месяцев назад
Explanations
1. 📊 Data science is a multi-disciplinary field using scientific methods to extract knowledge from structured, semi-structured, and unstructured data. The other options don't capture this comprehensive definition.
2. 📄 Data is defined as representation of facts, concepts, or instructions in a formalized manner suitable for communication and processing. The other options are too narrow or incorrect.
3. 📋 Information is processed data on which decisions and actions are based, having meaning to the recipient. The other options describe unprocessed data or unrelated concepts.
4. 📥 The data processing cycle consists of input (preparing data), processing (transforming data), and output (collecting results). The other options don't represent the complete cycle.
5. 🔢 Integer (int) data type is used to store whole numbers. Booleans store true/false, characters store single letters, and floats store decimals.
6. ✅ Boolean data type represents values restricted to true or false. The other options describe different data types.
7. 📋 Structured data adheres to a pre-defined model with tabular format and relationships between rows/columns. The other options describe unstructured characteristics.
8. 📑 JSON and XML are forms of semi-structured data with tags/markers to separate elements. SQL databases and Excel are structured; audio is unstructured.
9. 📁 Unstructured data doesn't have a predefined model or organized manner, like audio, video, or text-heavy content. The other options describe structured data.
10. 🔍 Metadata is data about data, providing additional information about a specific dataset (e.g., when/where photos were taken). The other options misdefine metadata.
11. 🔄 Data acquisition is gathering, filtering, and cleaning data before storage. The other options don't describe this initial value chain step.
12. 🎯 Data analysis makes raw data amenable for decision-making through exploration, transformation, and modeling. The other options don't capture this purpose.
13. 📊 Data curation is active management over data's lifecycle to ensure quality, including creation, selection, and validation. The other options misrepresent curation.
14. 📦 Data storage provides persistence and scalable management satisfying fast access needs. The other options contradict storage goals.
15. 💰 Data usage covers data-driven business activities requiring data access, analysis, and integration with business processes. The other options are too narrow.
16. 📏 Big data is characterized by Volume (large amounts), Velocity (streaming), and Variety (diverse forms), plus Veracity. The other V combinations are incorrect.
17. 📦 Volume refers to large amounts of data (Zettabytes/massive datasets) that exceed traditional processing capacity. The other options misinterpret volume.
18. 🌊 Velocity means data is live streaming or in motion, arriving at high speed. The other options don't relate to data speed.
19. 📊 Variety indicates data comes in many forms (structured, unstructured, semi-structured) from diverse sources. The other options aren't data-related.
20. 🔍 Veracity addresses whether we can trust the data and how accurate it is. The other options don't relate to data trustworthiness.
21. 🔗 Clustered computing combines resources of many smaller machines to handle big data needs. The other options describe single-machine setups.
22. 💾 Resource pooling combines storage, CPU, and memory from multiple machines for processing large datasets. The other options describe limitations, not benefits.
23. ⚙️ High availability provides fault tolerance preventing hardware/software failures from affecting data access. The other options don't address reliability.
24. ➕ Scalability allows adding machines horizontally to react to changing resource requirements. The other options describe limitations.
25. 🔧 YARN stands for Yet Another Resource Negotiator, handling cluster membership and resource allocation. The other acronyms are incorrect.
26. 💾 HDFS is Hadoop Distributed File System for storing data across clusters. The other acronyms are made up.
27. ⚡️ Spark performs in-memory data processing for faster big data analytics. The other options misinterpret Spark's purpose.
28. 🔍 PIG is a query-based data processing service in Hadoop ecosystem. It's not an animal, game, or app.
29. 📊 HBase is a NoSQL database for big data scenarios. SQL databases lack scalability for big data.
30. 🤖 Mahout and Spark MLLib provide machine learning algorithm libraries. The other uses are incorrect.
31. 🔎 Solr and Lucene provide searching and indexing capabilities in Hadoop. The other functions are incorrect.
32. 🔧 Zookeeper manages clusters in Hadoop ecosystem. It's not related to animals, games, or email.
33. 📅 Oozie handles job scheduling in Hadoop workflows. The other scheduling types are incorrect.
34. 📊 MapReduce is programming-based data processing in Hadoop. It's not for ...
for full explanation DM 👉@PaRatest123 on telegram