Добавил:
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

ENGLISH FOR IT. Учебное пособие

.pdf
Скачиваний:
0
Добавлен:
07.09.2026
Размер:
2 Мб
Скачать
Unit 8. Big Data
121
DEFINITIONS
TERMS
1
This term refers to the amount of data that is larger than ever before and constantly growing (6 letters).
_ _ _ _ _ _
2
This term refers to the increasing rate at which data is gen­erated and processed (8).
_ _ _ _ _ _ _ _
3
Quantitative data that consists of numbers and values and has a pre-defined data model (10, 4).
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _
4
Qualitative data that does not have a pre-defined data model, so it is best managed in non-relational (NoSQL) databases (12, 4).
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _
5
Data which provides additional information about a spe­cific set of data (8).
_ _ _ _ _ _ _ _
6
A mathematical formula or a set of instructions that we provide to the computer which describes how to process the given data in order to obtain needed information (9).
_ _ _ _ _ _ _ _ _
7
A digital collection of data that is organized in a specific way (8).
_ _ _ _ _ _ _ _
8
It stores data in tables that are related to each other (10, 8).
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _
9
A logical structure made up of rows and columns (5).
_ _ _ _ _
10
A data management system which aggregates large vol­umes of data from multiple sources into a single reposi­tory of highly structured and unified historical data (4, 9).
_ _ _ _ _ _ _ _ _ _ _ _ _
11
A repository that stores a huge amount of raw data in its original format. It employs a flat architecture which al­lows you to store raw data at any scale without the need to structure it first (4, 4).
_ _ _ _ _ _ _ _
12
A set of tools and processes used to automate the move­ment and transformation of data between a source system and a target repository (4, 8).
_ _ _ _ _ _ _ _ _ _ _ _
13
An approach to data analysis that involves building and adapting models, which allow programs to "learn" through experience and to improve their ability to make predic­tions (7, 8).
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ 14
Programming language used to manage structured data in relational databases (3).
_ _ _
15
The logical representation of a database, which shows how the data is stored logically in the entire database (6).
_ _ _ _ _ _
Unit 8. Big Data
122
USEFUL GRAMMAR
Grammar to study: Gerund. Verb patterns
7. Choose the correct verb to complete the sentence and use it as a Ger-
und. Translate the sentences into Russian.
benefit, make, handle, process, develop, place, work,
determine, discover, analyze, keep up, store, act, integrate, derive
1. _____ with big data technology is an ongoing challenge.
2. One of the biggest obstacles to _____ from your investment in big data is
a skills shortage.
3. _____ meaning in your data is not always straightforward.
4. Nearly every department in a company can utilize findings from big data
analysis, but _____ its clutter and noise can pose problems.
5. Many companies, such as Alphabet and Meta (formerly Facebook), use
big data to generate ad revenue by _____ targeted ads to users on social media and those surfing the web.
6. Over the period of time, talent in computer science has achieved greater
success in _____ techniques for _____ with such kind of data (where the format is well known in advance) and also _____ value out of it.
7. Size of data plays a very crucial role in _____ value out of data.
8. The final step is _____ and _____ on big data − otherwise, the investment
won’t be worth it.
9. _____ disparate data sources and _____ data accessible for business users
is complex, but vital, if you hope to realize any value from your big data.
10. Without the appropriate solutions for _____ and _____, it would be im-
possible to mine for insights.
8. Use the verb in brackets as a Gerund or an Infinitive (full or bare).
Translate into Russian.
1. This includes _____ (use) tools _____ (create) data visualizations like
charts, graphs, and dashboards.
2. Beyond _____ (explore) the data itself, it’s also critical _____ (communi-
cate) and share insights across the business in a way that everyone can _____ (understand).
Unit 8. Big Data
123
3. The enterprises that take steps now and make significant progress to-
ward _____ (implement) big data stand _____ (come) as winners in the future.
4. _____ (develop) a solid data strategy starts with _____ (understand) what
you want _____ (achieve), _____ (identify) specific use cases, and the data you currently have available _____ (use).
5. From the speed at which it’s created to the amount of time needed _____
(analyze) it, everything about big data is fast. Some have described it as _____ (try) _____ (drink) from a fire hose.
6. Analytical systems are more sophisticated than their operational counter-
parts, capable of _____ (handle) complex data analysis and _____ (provide) busi­nesses with decision-making insights.
7. Apache Spark is an open-source analytics engine used for _____ (process)
large-scale data sets on single-node machines or clusters.
8. Able to process over a million tuples per second per node, Apache
Storm’s open-source computation system specializes in _____ (process) distrib­uted, unstructured data in real time.
9. More and more companies begin _____ (move) their Enterprise Resource
Planning Systems (ERP) to the cloud.
10. By _____ (analyze) these indications of potential issues before the prob-
lems happen, organizations can _____ (deploy) maintenance more cost effec­tively and _____ (maximize) parts and equipment uptime.
11. _____ (leverage) this approach can _____ (help) _____ (increase) big
data capabilities and overall information architecture maturity in a more struc­tured and systematic way.
SPECIALIST READING
9. Read the article and fill the gaps with one of the following phrases:
1. Hence, ‘Volume’ is one characteristic which needs to be considered while
dealing with Big Data solutions
2. That starts with data pipelines that process using ETL (extract, trans-
form, load)
3. The flow of data is massive and continuous
4. They are just used differently
5. They describe some of the characteristics that make big data different
from other data processing
Unit 8. Big Data
124
6. This is typically done separately in a processing system
7. That could be anywhere from multiple storage types to including rela-
tional databases and/or data warehouses
8. Nowadays, data in the form of emails, photos, videos, monitoring devices,
PDFs, audio, etc. are also being considered in the analysis applications
9. This allows for parallel computations to occur on larger data sets
10. This is considered schema-on-read
If you have data that’s too large to store on a single computer, data that can
rapidly grow, or would be too difficult or take too long to process using any tra­ditional methods, you have Big Data. The inflow of data can also be unpredictable as the datasets grow in diversity that may be structured or unstructured. Big Data, whether by complexity or sheer volume, is much more difficult to process with standard methods.
In 2001, Gartner’s Doug Laney first presented what became known as the “three Vs of big data”. (1) _____:
Volume – the name Big Data itself is related to a size which is enormous. Size of data plays a very crucial role in determining value out of data. Also, whether a particular data can actually be considered as a Big Data or not, is de­pendent upon the volume of data. (2) _____.
Velocity – the term ‘velocity’ refers to the speed of generation of data. How fast the data is generated and processed to meet the demands, determines real potential in the data. Big Data Velocity deals with the speed at which data flows in from sources like business processes, application logs, networks, and social media sites, sensors, Mobile devices, etc. (3) _____.
Variety – variety refers to heterogeneous sources and the nature of data, both structured and unstructured. During earlier days, spreadsheets and data­bases were the only sources of data considered by most of the applications. (4) _____. This variety of unstructured data poses certain issues for storage, mining and analyzing data.
Because data would be too large to store and process, Big Data is handled differently in storage. Instead of a database on a computer, one storage place for Big Data is a Data Warehouse.
To store and process large quantities of data, a data warehouse is the central point of data storage. It is a system that allows data to flow into a single source, supports analytics, data mining, machine learning, and so on. Although its goal
Unit 8. Big Data
125
is to store data in one location, though not on a single computer, another objective is to process through the data. (5) _____. The warehouses collect the data and format it in a usable way. Wherever the data comes from, it starts as raw data which can be processed into something usable.
A data warehouse deals with large volumes of data which can be old or new data. However, the data it stores must be structured. The data may then be orga­nized into schemas for later analysis.
Once the data is stored, it can be processed. (6) _____. Processing needs to be done separately because the amount of data, or even its complexity, put heavy demands on the underlying computing infrastructure. In many instances, the pro­cessing is performed in the cloud.
When data is unstructured or even unrelated, data lakes can be used. Data lakes store data without a defined schema, so the data would be unable to store in
a relational database. But that doesn’t mean data lakes can’t process data. Data
lakes can create various visualizations, real-time analytics, and so on with the need to structure the data first. Like a warehouse, it is a centralized storage place that can be fed by multiple sources. The raw data is stored for later processing, but instead of structuring and adding schemas, data lakes can simply store the data without any grooming.
We just mentioned that you don’t need to create any schemas to fit data into,
as it does not need to be structured. Schemas, however, are still a part of data lakes. (7) _____. For example, data warehouses clean and filter data into sche­mas, which are designed before preprocessing the data. This is considered schema-on-write. With a data lake, you don’t design the schema ahead of time. Instead, the schema is created while the data is being analyzed. (8) _____.
To store the data, different storage devices could be used. Cloud storage could be considered. If databases are required, data lakes use unstructured or non­relational models such as NoSQL databases. Data lakes can also exist on devices such as a Hadoop cluster. A Hadoop cluster uses a series of computers, called nodes, which are networked together (9) _____.
There is always the possibility that the system you select for your data lake may not be enough. That is why data lakes can be combined with multiple sys­tems using a distributed architecture. In this case, the data lake would be the cen­tralized point for the storage and processing but can branch with other platforms. (10) _____. Although data lakes allow the inflow of data to remain in its raw form
Unit 8. Big Data
126
for later processing, you may also choose to preprocess with different data mining tools or data preparation software [21].
10. Read the article again and say whether the following statements are
true or false.
STATEMENTS
T/F
1
Big Data is difficult to process with standard methods because of its complexity and size.
2
The “three Vs of big data” describe different types of big data.
3
We consider a particular data as Big Data dependent upon the volume of data. 4
Velocity refers to the data inflow rate.
5
Nowadays spreadsheets and databases are the only sources of data con­sidered by most of the applications.
6
Big Data is stored the same way as traditional data.
7
The goal of a data warehouse is to store data on a single computer.
8 A data warehouse deals with structured data.
9 Data lakes can’t process data.
10
With a data lake, the schema is created while the data is being analyzed.
11. Answer the following questions according to the article.
1. What differs big data from traditional data?
2. Why is Big Data handled differently in storage?
3. What does ‘variety’ refer to?
4. What is the purpose of a data warehouse?
5. What kind of data do data lakes store?
6. Can data lakes create real-time analytics?
7. What do data warehouses and data lakes have in common?
8. Can data lakes branch with relational databases and data warehouses?
9. What do data lakes use if databases are required?
10. What does schema-on-read mean?
12. Summarize the article from exercise 9. Be guided with the tips on
page 201.
VOCABULARY IN USE
13. Follow the link https://rutube.ru/audio/355128e7cb0af21c8cc819b-
05260341c/ to listen to the article. Then read the article and fill the gaps.
Unit 8. Big Data
127
BIG DATA
Big data refers to massive, complex data sets (either (1) _____, semi-struc­tured or unstructured) that are rapidly generated and transmitted from a wide (2) _____of sources.
These days, data is constantly generated anytime we open an app, search Google or simply travel place to place with our mobile (3) _____. The result? Massive collections of (4) _____ information that companies and organizations manage, (5) _____, visualize and analyze.
Traditional data tools aren’t equipped to handle this kind of complexity and
(6) _____, which has led to a slew of specialized big data software (7) _____ and architecture solutions designed to manage the load.
Big data platforms are specially designed to handle huge volumes of data that come into the system at high (8) _____ and wide varieties. These big data platforms usually consist of varying servers, (9) _____and business intelligence tools that allow data scientists to manipulate data to find trends and patterns.
Variety, Volume and Velocity make up the (10) _____ of big data. Let’s take a closer look at each attribute.
Volume
Big data is (11) _____. While traditional data is measured in familiar sizes like megabytes, gigabytes and terabytes, big data is stored in petabytes and zet­tabytes.
To grasp the enormity of the (12) _____ in scale, consider this comparison from the Berkeley School of Information: One gigabyte is the equivalent of a seven minute video in HD, while a single zettabyte is equal to 250 billion DVDs.
Big data provides the (13) _____ handling this kind of data. Without the appropriate solutions for storing and (14) _____, it would be impossible (15) _____ for insights.
Velocity
From the speed at which it’s created to the amount of time needed
(16) _____ it, everything about (17) _____ is fast. Some have described it as try­ing to drink from a fire hose.
Companies and organizations must have the capabilities to harness this data and (18) _____ insights from it in real-time, otherwise it’s not very useful. Real- time processing allows (19) _____ to act quickly, giving them a leg up on the competition.
Unit 8. Big Data
128
Variety
Roughly 80 to 90 percent of all big data is (20) _____, meaning it does not fit easily into a straightforward, traditional model. Everything from emails and videos to scientific and meteorological data can constitute a big data stream, each with their own unique attributes [22].
14. Read the text and fill the gaps with one of the given words:
data lakes, to process, on write, Hadoop, velocity, volume,
analysis, big data, data pipelines, storing data, warehouse,
on read, ETL, schema, structured, raw data
Big data is about more than just data (1) _____. Two other characteristics of "the 3 V's" are variety and (2) _____. Big data refers to data that is so large, fast or complex that it is difficult or impossible (3) _____ using traditional methods.
Modern computing systems provide the speed, power and flexibility needed to quickly access massive amounts and types of (4) _____. Along with reliable access, companies also need methods for integrating the data, building (5) _____, ensuring data quality, providing data governance and storage, and preparing the data for (6) _____. Some big data may be stored in a traditional (7) _____ – but there are also flexible, low-cost options for (8) _____ and handling big data via cloud solutions, data lakes, data pipelines and (9) _____.
A data warehouse is a tool that has become synonymous with extract, trans­form and load ((10) _____) processes. At a high level, data warehouses store vast amounts of (11) _____ data in highly regimented ways. They require that a pre­defined (12) _____ exists before loading the data, Put differently, the schema in a data warehouse is defined “(13) _____”. For decades, data warehouses have handled even large volumes of structured data exceptionally well. It’s unreason- able, however, to expect those same data warehouses to efficiently process fun-
damentally different data volumes, speeds and types, so we’ve seen the rise in
popularity of (14) _____.
The data lake is fundamentally different in the following regard. As David
Loshin writes, “The idea of the data lake is to provide a resting place for
(15) _____ in its native format until it’s needed.” Data lies dormant unless and until someone or something needs it. For this very reason, a data lake schema is defined “(16) _____. Put differently, a data lake still requires a schema. How­ever, that schema is not predefined [23].
Unit 8. Big Data
129
SPEAKING
Structuring a presentation
LANGUAGE WORK
15. Complete the sentences with the words from the box.
after – all – areas – divided – finally – start – then – third
a) I’ll be talking to you today about the Big Data Services we offer.
I’ll _____ (1) by describing our big data consulting services. _____ (2) I’ll go on
to show you some case studies. _____ (3), I’ll discuss how you can choose the best solutions to help our clients become a truly digital business.
b) I’ve _____ (4) my talk into three main parts. First of _____ (5), I’ll tell
you something about the history of our company. _____ (6) that I’ll describe how
the company is structured and finally, I’ll give you some details about our range
of products and services.
c) I’d like to update you on what we’ve been working on over the last year. I’ll focus on three main _____ (7): first, AI-based products and services; second, industry-specific solutions. And _____ (8), advanced security measures.
16. Complete the sentences with the prepositions from the box.
into − on (2) − on to to (2) with about for after
1. 'My presentation is divided ___ three main sections.'
2. 'In my presentation I’ll focus ___ three major issues.'
3. 'I’m going to concentrate ___ the functions of big data companies.'
4. 'I’d now like to move ___ data security and privacy.'
5. 'I'd also like to draw your attention ___ advantages of collaborating with
big data consulting firms.'
6. 'As you remember, we are concerned ___ multi-cloud management solu-
tions.'
7. 'That brings me ___ the end of my presentation.'
8. 'That’s all I have to say ___ data management.'
9. 'Thank you ___ listening.'
10. 'I’d be grateful if you could ask your questions ___ the presentation.'
17. Match the elements from each column to make “signpost” sentences.
Unit 8. Big Data
130
Before I move on to my next point,
come back to
next question. This brings
the issue
point, which is data lake.
This leads
let me go
this question later.
Let’s now turn to
we were discussing
our new solutions that help illustrate the data most com­prehensively.
As I mentioned
to the next
a brief overview of our big data services.
I’d like to
before, I’d like to give you
earlier.
Let’s go back to what
us directly to my
through the main issues once more.
As I said earlier,
I’ll be focusing on
of big data mining.
18. Prepare and give a talk about one of the Data Science tools (BigML,
Apache Spark, SAS Excel Tableau etc.). Be guided with the tips on page 197 and on the site https://www.eureka.org.nz/how-to-structure-your-talk.
WRITING
Writing an email cover letter
19. Put the phrases in the correct group.
1. I’ve had this job for the past two years.
2. The bulk of my work involves …
3. I can work across a range of platforms.
4. I am very comfortable using analytics.
5. I have always worked in IT.
6. I have the skills to push business forward through creativity and
innovation.
7. My new website was responsive, lightning fast, and included the latest e-
commerce features.
8. I’ve been programming websites and using CSS to create user-friendly
experiences since I was in nine form, so it’s long been a passion of mine.
Talking about work experience
Talking about transferable skills