Методология научного творчества = Methodogy of scientific research. Учебно-методическое пособие
.pdf3.the information request should be translated to the information retrieval language[11]
- highlight keywords;
- define the language search framework;
- determine the chronological scope of the search.
4.To specify, whether there is no ready bibliography on a subject
5.Ifthereisaready-madebibliography,supplementitwithnew literature. (If there is no finished bibliography, highlight a retrospective search for information on catalogs, card indexes and databases of libraries).
6.Or find a textbook on the research topic (read the article in Wikipedia) and use its bibliographic list, finding sources from it and their bibliographic lists, respectively.
7.If necessary, translate keywords into a foreign language and search for foreign literature.
To characterize the properties of information in the Global Network, it is customary to define two aspects: the relevance and pertinence. In order to perform a search on the Web, you need to formulate a query.
Relevance - this is the correspondence of the information received to the search query (when searching for "cats" the search engine will give out all the sites, in the text of which cats were mentioned).
Pertinence is the correspondence of the information received to the user's needs (if the user was looking for a "cat" device to climb poles or trees (in Russian these words are the same), a possible answer to his request will be "buried" in a huge amount of information about cats (animals)).
The concept of relevance is already narrower than pertinence A document issued by a search engine may be relevant to the request, but not satisfy the information need of the user. The reason for this is the ambiguity and insufficiency of the natural language [9].
11
If the user can not directly influence the work of search engines, then the quality of the search request is entirely in his competence.
Search Engine Strategies
Information retrieval system means (IRS) information service, designed to find in a set of documents that contain information relevant to the information request.
IRS consists of three interrelated elements:
-documents containing certain information;
-information retrieval language (a list of terms and their application);
-special media and their corresponding devices used for
information retrieval.
Depending on the object of storage and the type of request, there are two types of information retrieval: documentary and factographic.
The third one – is the information-logical system which responsetoquerieswithnoexplicitanswerintheinformationbase. Extralinguistic knowledge base and information generated algorithmically from the already available (documentary or factographic)helptogettheanswer.Thisnewinformationiseither issuedasaresponsetoarequestoradditionallyusedforsearching.
The main search facilities on the Internet are information retrieval systems (IRS) of three types: classification, vocabulary, and subject.
1. Classification IRS - the hierarchical organization of information. First, the team of authors develops a classifier, then another team (systematizes) places documents and sections on its headings and sections. Examples of classification IPS: Excite, Look Smart Yellow Web, "Constellation Internet", etc. Disadvantages of classification IRS: Classifiers, created by different teams in different countries vary greatly (mentality,
12
cultural differences, difficulties in interpreting foreign materials, the possibility of classifying documents under different headings simultaneously).
2.VocabularyIRS - use a database builtfromwords found in Internet documents. In such a database, every word contains a list of documents from which it is taken. Dictionary IPS of the Internet are AltaVista, Rambler, Yandex, Aport.
3.Subject IRS - lists of web resources containing the necessary information and links to related sites
Search engines use different algorithms to sort the results:
1)the number of query words in the text of the document;
2)headers with these words;
3)position of words in the document;
4)thepercentageofrelevantwordsinthetotalnumberofwords in the document;
5)time limit - how long the page is in the database of the search server.
6)citation index - how many links to this page leads from other pages registered in the database search engine [12].
Search, as well as searching in a regular library, can be divided into simple and advanced. Each of them has a certain set of receptions (cited by:[13])
Simple Search Techniques
1.Search for a group of words
The word "ecology" or "situation" is too general and will
give a large number of references pertaining to different topics when looking for separately, but it is unlikely that among them the concept of "ecological situation" will occur. In connection with these, it is recommended to add keywords that narrow the search.
13
2.Search for word forms
In most cases, the search engine searches for all the word
variations by default. In order not to go through all the word forms from the query when searching, you can use an exclamation point: the query "! Fridge" will find pages with this word without regard for single-root words and cases.
3.The role of capital letters
If the keyword is entered as a query with a capital letter,
there will be no pages in the response, with this word from the lowercase letter. In this regard, it is recommended to use capital letters only in the names of your own. For example, "city of Moscow".
4.The value of wildcard characters
The symbol "*" can be used instead of any number of any
characters up to the end of the word.
5.Accounting for reserved words
"Reserved words" (stop words) are words that are not
considered when searching (less than 4 letters). These are pretexts, unions, etc. The query "in Germany" - found only documents that include the word "Germany" or its variants.
6.Means of contextual search
If the keywords are quoted, the IPS will find all the
documents in which this quote is present completely.
Advanced search methods (quoted in [13])
For faster and more successful search in search engines, different logical operators are used in conjunction with keywords. Therulesforcompilingcomplexqueriesononesearchenginemay differ from those for the other, but in any case, the following basic operators will be used:
14
1. Operator AND
Combining two or more words so that they are all present in the document you are looking for. Often, instead of AND, use & or +. Example: on request, the lawyer and the program will find documents that contain both words.
2. Operator OR
The search is conducted according to any of the words of thegroup.Forexample:onrequesteducationORtrainingwillfind documents containing only ONE word.
3. Boolean Brackets
Applied when you need to specify the order of the logical operators. For example, on request, Putin OR (Vladimir I Putin) will find documents containing the words Putin or Vladimir I Putin.
4. Operator NOT
It is used when from the search results it is necessary to exclude any keyword, for example, on the request of Shakespeare NOT Hamlet, plays of the great bard will be found, except for Hamlet.
5. Operator NEAR
Search by distance. It says that words should be in the immediate environment of the keyword. The syntax of such a query is different for different search engines.
Table 3 The most popular search engines[14], [15], [16], [17]
1. |
The search engine giant holds the first place in |
||||
|
|
search |
|
|
|
2. |
Bing |
Microsoft’s attempt to challenge Google |
|
||
3. |
Yahoo |
holds the third place in search |
|
|
|
4. |
Ask.com |
is based on a question/answer format where most |
|||
|
|
questions are answered by other users or are in the |
|||
|
|
form of polls |
|
|
|
5. |
AOL.com |
the network includes many popular websites like |
|||
|
|
engadget.com, |
techchrunch.com, |
and |
the |
|
|
huffingtonpost.com. |
|
|
|
15
6. Baidu |
|
the most popular search engine in China |
7. |
|
Computational Knowledge Engine which can give |
WolframAlpha |
you facts and data on a number of topics. It can do |
|
|
|
all sorts of calculations, for example, if you enter |
|
|
"mortgage2000"asinputitwillcalculateyourloan |
|
|
amount, interest paid etc. based on a number of |
|
|
assumptions. |
|
|
The distinguishing feature of this search is: you |
|
|
will not see any references to posts in social |
|
|
networks or articles of the yellow press, only |
|
|
specific figures and proven facts in the form of a |
|
|
single document. |
8. DuckDuckGo |
It has a clean interface, it does not track users, it is |
|
|
|
not fully loaded with ads and has a number of very |
|
|
nice features (only one page of results, you can |
|
|
search directly other websites etc). |
9. |
Internet |
the internet archive search engine. You can use it |
Archive |
|
to find out how a website looked since 1996. It is |
|
|
the very useful tool if you want to trace the history |
|
|
of a domain and examine how it has changed over |
|
|
the years. |
10. ChaCha.com |
It is similar to ask.com where users can ask or |
|
|
|
answer a particular question. They also have a |
|
|
number of quizzes that can help you decide on a |
|
|
number of topics. |
11 Rambler |
The search index is updated daily, so it's easy to |
|
|
|
find the latest information; |
|
|
Since 2011, by agreement with Yandex, uses his |
|
|
search algorithm and is no longer an independent |
|
|
search engine. |
12 |
|
For the ranking of sites used about 250 factors, |
search.mail.ru |
including the behavioral factor. |
|
13 Yandex.ru |
Search engine uses the algorithm of personalized |
|
|
|
search and geo-dependency queries, depending on |
|
|
the region of the site and the user; |
|
|
Yandex has a system of hints, bugs are fixed; |
16
14 |
Teoma |
Subject-Specific Popularity, Hyperlink-Induced |
|
|
Topic Search, Refinement of the query in the form |
|
|
of several key phrases on the subject of the request |
|
|
and links to pages on the query subject prepared by |
|
|
a team of experts and enthusiasts |
15 |
LookSmart |
Search in own catalog |
16 |
Freeserve.net |
Search results for the base of the British version of |
|
|
Overture |
17 not Evil |
A system that searches the anonymous Tor |
|
http://hss3uro2h |
network. Looks for where Google, "Yandex" and |
|
sxfogfq.onion/ |
other search engines are closed in principle. |
|
18 |
YaCy |
A decentralized search engine that operates on the |
|
|
principle of P2P networks. Each computer on |
|
|
which the main program module is installed scans |
|
|
the Internet independently, that is, it is analogous |
|
|
to the search robot. The results obtained are |
|
|
collected in a common database, which is used by |
|
|
all YaCy participants. |
19 |
Pipl |
System designed to find information about a |
|
|
particular person |
20 |
FindSounds |
Specialized search engine. Looks for different |
|
|
sounds (home, nature, cars, people, etc.) in open |
|
|
sources |
21 |
The Apache |
Free library for high-speed full-text search, |
Lucene |
|
|
22 |
Wikia |
Free open search engine and part of Wikia |
Search |
|
|
23 |
A special search engine for the world of scientific |
|
Scholar |
literature |
|
24 |
Спутник |
Was created at the request of the Russian |
|
|
leadership. Sputnik is the world's first state-owned |
|
|
search engine. This is both its plus and minus. |
25 |
|
a platform for academics to share research papers. |
Academia.edu |
The company's mission is to accelerate the world's |
|
|
|
research. Academics use Academia.edu to share |
|
|
their research, monitor deep analytics around the |
17
|
|
impact of their research, and track the research of |
||
|
|
academics they follow. |
|
|
26 |
Mendeley |
Mendeley Data is a secure cloud-based repository |
||
|
|
where you can store your data, ensuring it is easy |
||
|
|
to share, access and cite, wherever you are. |
||
27 |
Dogpile.com |
One of the oldest meta-search engines |
|
|
28 |
Zoo.com |
(formerly Metacrawler.com). Another oldest and |
||
|
|
well-known in the English-language Internet meta- |
||
|
|
search engine. Uses the results of major search |
||
|
|
engines. |
|
|
29 metabot.ru |
Russian multi meta search engine |
|
||
30 |
|
large-scale, |
high-performance, |
real-time |
gigablast.com |
information retrieval technology and services for |
|||
|
|
partner sites. The company offers a variety of |
||
|
|
features including topic generation and the ability |
||
|
|
to index multiple document formats. |
|
|
While more people use the Internet than ever before for their research, this is notwithoutitstroubles. TheInternetcontains valuable information, but it also contains information that has not been well-researched [16].
Someofthescientificresearchinformationavailableonthe internet are[18]:
•Details about various scientific and non-scientific topics.
•Titles and other relevant information of article published in various journals, possibly,
•From past one decade or so (full article will not be available).
•Preprint of papers submitted by researchers in certain websites.
•Information about scientific meetings to be held.
•Contact details for other researchers.
•Databases of reference material.
•Places where one can discuss topics and ask for help.
In general, academic research that has been commercially published is not freely available on the internet.
18
STEP 3: Introduction: CRITERIA OF NOVELTY AND
ACTUALITY
Scientific novelty is the criterion of scientific research, determining the degree of transformation, supplementation, and concretization of scientific data [19]. The criterion of novelty characterizes the content side of the result, new theoretical positions and practical recommendations that were not previously known and were not fixed in science and practice.
There are 3 levels of scientific novelty: [20]
A)the transformation of known data, a fundamental change in it
B)the expansion and addition of known data without changing essence
C)the refinement, specification of known data, dissemination of known results to a new class of objects or systems
Kind of novelty [14]
•Theoretical novelty (concept, hypothesis, terminology, etc.)
•Practical (rule, proposal, recommendation, means, requirement, methodological system, etc.).
Thelevelofnovelty-characterizestheplaceofknowledge acquired in a number of known and their continuity. It is evaluated using the level of specification, addition, and transformation.
As scientific novelty can act [20]:
•Knowledge
o the problem has been posed and considered for the first time or
onew formulation of known problems or tasks (for example, assumptions are removed, new conditions are accepted);
•Method:
oa new method of solution;
o a new application of a known solution or method;
19
•tool (new improved instruments to obtain new results from the same conditions)
•implementation (new or improved criteria, indicators)
It is no obligation to all available elements of novelty to be presented in the work.
The criterion of relevance indicates the necessity and timeliness of studying and solving the problem for the further development of the theory and practice of education and upbringing,characterizesthecontradictionsarisingbetweenpublic needs (the demand for scientific ideas and practical recommendations) and the means of their satisfaction. The criterion of relevance is dynamic, depends on time, specific conditions and specific circumstances.[14]
Once thetopic ofthestudy is determined, anditisrelevant, it is necessary to identify the main goal of the study. The goal follows from the urgency and is usually expressed in the form of "study ...", "determine ...", etc.
Unlike the goal, which usually has one study, there are several problems. They are intermediate stages on the way to achieving the goal, fulfill the role of the research plan and are necessarily reflected in the conclusions that sum up these stages.
20
