Saturday, March 28, 2020
7 Proofreading Steps
7 Proofreading Steps 7 Proofreading Steps 7 Proofreading Steps By Mark Nichol Proofreading is the last line of defense for quality control in print and online publishing. Be sure to conduct a thorough proofread of all documents before they are printed for distribution and of all Web pages before they go live, using these guidelines. But before you proof, you must edit. (This post explains the difference between the two processes.) Thereââ¬â¢s no use expending time and effort to check for minor typographical errors until the editing stage is complete. Review for proper organization, appropriate tone, and grammar, syntax, usage, and style before the document is laid out. Stakeholders should read the edited version before layout and submit requests for revisions during the editing stage. If anyone other than the editorial staff must see the proof, remind him or her that only minor changes should be made at this point. 1. Use a Checklist Create a list of important things to check for, such as problem areas like agreement of nouns and verbs and of pronouns and antecedents, and number style. 2. Fact-Check Double-check facts, figures, and proper names. If information remains to be inserted at the last minute, highlight the omission prominently so that no one forgets to do so. 3. Spell-Check Before proofreading a printout, spell-check the electronic version to find misspellings, as well as errors you or a colleague make frequently, such as omitting a closing parenthesis or quotation mark. 4. Read Aloud Reading text during the proof stage improves your chances of noticing errors, especially missing (ââ¬Å"a summary the report followsâ⬠) or repeated (ââ¬Å"a summary of the the report followsâ⬠) words. 5. Focus on One Line at a Time When proofing print documents, use another piece of paper or a ruler to cover the text following the line you are proofreading, shifting the paper down as you go along. This technique helps you keep your place and discourages you from reading too quickly and missing subtle errors. 6. Attend to Format Proofreading isnââ¬â¢t just about reviewing the text. Make sure that the document design adheres to established specifications. Check page numbering, column alignment, relative fonts, sizes, and other features of standard elements such as headlines, subheadings, captions, and footnotes. Inspect each type of feature within categories, such as looking at every headline, then every caption, and so on. 7. Proof Again Once revisions have been made, proofread the document again with the same thoroughness, rather than simply spot-checking the changes. An insertion or deletion may have thrown off the line count, for example. Want to improve your English in five minutes a day? Get a subscription and start receiving our writing tips and exercises daily! Keep learning! Browse the Writing Basics category, check our popular posts, or choose a related post below:Direct and Indirect Objects41 Words That Are Better Than GoodIs "Number" Singular or Plural?
Saturday, March 7, 2020
Aaron Burr - Biography and the Duel with Hamilton
Aaron Burr - Biography and the Duel with Hamilton Aaron Burr is mostly remembered for a single violent act, the fatal shooting of Alexander Hamilton in their famous duel in New Jersey on July 11, 1804. But Burr was also involved in a number of other controversial episodes, including one of the most disputed elections in American history and a peculiar expedition to the western territories that resulted in Burr being tried for treason. Burr is a puzzling figure in history. He has often been portrayed as a scoundrel, a political manipulator, and a notorious womanizer. Yet during his long life Burr had many followers who considered him a brilliant thinker and a gifted politician. His considerable skills allowed him to prosper in a law practice, win a seat in the U.S. Senate, and nearly attain the presidency in a startling feat of deft political gamesmanship. After 200 years, Burrââ¬â¢s complicated life remains contradictory. Was he a villain, or simply a misunderstood victim of hardball politics? Early Life of Aaron Burr Burr was born in Newark, New Jersey, on February 6, 1756. His grandfather was Jonathan Edwards, a famous theologian of the colonial period, and his father was a minister. Young Aaron was precocious, and entered the College of New Jersey (present day Princeton University) at the age of 13. In the family tradition, Burr studied theology before becoming more interested in the study of law. Aaron Burr in the Revolutionary War When the American Revolution broke out, the young Burr obtained a letter of introduction to George Washington, and requested an officers commission in the Continental Army. Washington turned him down, but Burr enlisted in the Army anyway, and served with some distinction in a military expedition to Quebec, Canada. Burr did later serve on Washingtonââ¬â¢s staff. He was charming and intelligent, but clashed with Washingtonââ¬â¢s more reserved style. In ill health, Burr resigned his commission as a colonel in 1779, before the end of the Revolutionary War. He then turned his full attention to the study of the law. Burrs Personal Life As a young officer Burr began a romantic affair in 1777 with Theodosia Prevost, who was 10à years older than Burr and also married to a British officer. When her husband died in 1781, Burr married Theodosia. In 1783 they had a daughter, also named Theodosia, to whom Burr was very devoted. Burrââ¬â¢s wife died in 1794. Accusations always swirled that he was involved with a number of other women during his marriage. Early Political Career Burr began his law practice in Albany, New York before moving to New York City to practice law in 1783. He prospered in the city, and established numerous connections that would prove useful in his political career. In the 1790s Burr advanced in New York politics. During this period of tension between the ruling Federalists and the Jeffersonian Republicans, Burr tended not to align himself too much with either side. He was thus able to present himself as something of a compromise candidate. In 1791, Burr had won a seat in the U.S. Senate by defeating Philip Schuyler, a prominent New Yorker who happened to be the father in law of Alexander Hamilton. Burr and Hamilton had already been adversaries, but Burrââ¬â¢s victory in that election caused Hamilton to hate him. As a senator, Burr generally opposed the programs of Hamilton, who was serving as secretary of the treasury. Burrs Controversial Role in the Deadlocked Election of 1800 Burr was the running mate of Thomas Jefferson in the presidential election of 1800. Jeffersonââ¬â¢s opponent was the incumbent president, John Adams. When the electoral vote produced a deadlock, the election had to be decided in the House of Representatives. In the prolonged balloting, Burrà utilized his considerable political skills and nearly pulled off the feat of bypassing Jefferson and gathering enough votes to win the presidency for himself. Jefferson finally won after days of balloting. And in accordance with the Constitution at the time, Jefferson became president and Burr became vice president. Jefferson thus had a vice president he didnââ¬â¢t trust, and he gave Burr virtually nothing to do in the job. Following the crisis, the Constitution was amended so the scenario of the 1800 election could not occur again. Burr was not nominated to run with Jefferson again in 1804. Aaron Burr and the Duel With Alexander Hamilton Alexander Hamilton and Aaron Burr had been conducting a feud since Burrââ¬â¢s election to the Senate more than 10à years earlier, but Hamiltonââ¬â¢s attacks on Burr became more intense in early 1804. The bitterness reached its climax when Burr and Hamilton fought a duel. On the morning of July 11, 1804 the men rowed across the Hudson River from New York City to a dueling ground at Weehawken, New Jersey. Accounts of the actual duel have always differed, but the result was that both men fired their pistols. Hamiltonââ¬â¢s shot did not strike Burr. Burrs shot struck Hamilton in the torso, inflicting a fatal wound. Hamilton was brought back to New York City and died the next day. Aaron Burr was portrayed as a villain. He fled andà actually went into hiding for a time, as he feared being charged with murder. Burrs Expedition to the West The once-promising political career of Aaron Burr had been stalled while he served as vice president, and the duel with Hamilton effectively ended any chance he may have had for political redemption. In 1805 and 1806 Burr plotted with others to create an empire consisting of the Mississippi Valley, Mexico, and much of the American West. The bizarre plan had little chance for success, and Burr was charged with treason against the United States. At a trial in Richmond, Virginia, which was presided over by Chief Justice John Marshall, Burr was acquitted. While a free man, his career was in ruins, and he moved to Europe for several years. Burr eventually returned to New York City and worked at a modest law practice. His beloved daughter Theodosia was lost in a shipwreck in 1813, which further depressed him. In financial ruin, he died on September 14, 1836, at the age of 80, while living with a relative on Staten Island in New York City. Portrait of Aaron Burr courtesy of New York Public Library Digital Collections.
Wednesday, February 19, 2020
Ethical aspects of science Essay Example | Topics and Well Written Essays - 1000 words
Ethical aspects of science - Essay Example In an academic context, ethics needs to be considered as ââ¬Ëan area of study that deals with ideas about what is good and bad behaviourââ¬â¢. Ethics in the academic context is commonly considered to be a branch of philosophy that deals with what is right and what is wrong from a moral point of view. à In general, ethics needs to be considered as all the moral principles that may influence our decisions and correct our behaviour. It needs to be pointed out that these principles can include working, eating, communicating with other people, studying, and so forth. à These principles are meant to keep our own lives and lives of people around in the right order. That is why, since the ancient times people had been expected to follow the rules of ethics and to encourage others to do the same. However, it is necessary to keep in mind that modern world is a complicated place where everything changes fast. à Consequently, the need to adopt ethical theories to the new conditions o f life arises.à Technologies and science develop new ideas faster than ever, and one of the major concerns of science in a context of ethics is a field of biomedical research. à Dramatically fast development of biomedical technologies that happened during the last twenty years produced a huge amount of ethical issues. It is necessary to mention that there is a list of reasons explaining why adherence to ethical norms is so important in a field of research.à Firstly, the aims of any research are knowledge, avoidance of errors.
Tuesday, February 4, 2020
Organizational Analysis Essay Example | Topics and Well Written Essays - 1500 words
Organizational Analysis - Essay Example Human resource plays a very important role in the development and success of any organization so was the case with Wal-mart. Sam Walton from the start of this business was surrounded by the most creative and hardworking employees. The employees are still working with their complete dedication and interest to achieve the goal of the mission statement. There are many internal and external challenges faced by the Human resource of Wal-mart such as the employee turnover rate, less capable employees in the developing countries, world politics, economics, inflation, exchange rates, etc. However, Wal-mart successfully faced all the hurdles in its way and qualified to be considered the largest retailers chain in the world. But there is always a room for further improvements and achievements and to fill that gap Wal-mart should continuously come up with new and different ideas to remain dominant in the retailerââ¬â¢s world. Organizational Analysis of Wal-mart Today, the customers not only want to buy things that they want but they actually want to enjoy their shopping experience. Now customers want a lot of merchandize available under one roof with the satisfying services and lowest possible prices, friendly and pleasant shopping environment with free parking. Wal-mart promises to give all of this to its customers (Walton, 2012). Wal-mart is a super store which features maximum number of high quality merchandize with comparatively low prices and gives its customers an everlasting shopping experience. It serves more than 200 million customers per week (our story, 2012). It has retail stores, online services and mobile alerts operating in 27 countries under 69 different banners. The first Wal-mart store was open in 1962 in Rogers, Arkansas. Sam Waltonââ¬â¢s unparallel devotion to the company and the leadership skills lead the organization to where it is now standing. He was the man behind the success of the unique retail store. He believed in leadership through serv ice and customer satisfaction. The basic idea behind Wal-mart was to serve the customers with low prices and great service. The target market of Wal-mart is that segment of customers who want multiple things such as grocery, electronics, apparel, stationary, decorative, and every other thing under one roof. These customers want a pleasant buying experience and goods services and satisfaction along with low prices. Wal-mart is very successful in fulfilling its customersââ¬â¢ requirements and therefore it has started the online and mobile services as well considering the current market trends and intense competition. The customers who believe in saving and spending good lives are the real customers of Wal-mart. In 1960, the whole idea of retail stores was changes as the Wal-mart step in the world of retailers. By 1967 Wal-mart was able to own 24 stores with $12.7 million sale (history timeline, 2012). Later in 1980ââ¬â¢s the first Wal-mart supermarket was opened with general mer chandise. In 1987 the company installed the largest satellite communication system in the United States of America. In 1990ââ¬â¢s Wal-mart was marked as the most successful and the biggest retail store. By 2002, Wal-mart was among the 500 ranking of the Americaââ¬â¢s companies. In 2012, the company has celebrated its 50th anniversary with 2.2 million associates, 200 million customers and 10,000 stores in 27 countries. Mission Statement Wal-mart was made with the mission of
Monday, January 27, 2020
Partitioning Methods to Improve Obsolescence Forecasting
Partitioning Methods to Improve Obsolescence Forecasting Amol Kulkarni Abstract Clustering is an unsupervised classification of observations or data items into groups or clusters. The problem of clustering has been addressed by many researchers in various disciplines, which serves to reflect its usefulness as one of the steps in exploratory data analysis. This paper presents an overview of partitioning methods, with a goal of providing useful advice and references to identifying the optimal number of cluster and provide a basic introduction to cluster validation techniques. The aim of clustering methods carried out in this paper is to present useful information which would aid in forecasting obsolescence. INRODUCTION There have been more inventions recorded in the past thirty years than all the rest of recorded humanity, and this pace hastens every month. As a result, the product life cycle has been decreasing rapidly, and the life cycle of products no longer fit together with the life cycle of their components. This issue is termed as obsolescence, wherein a component can no longer be obtained from its original manufacturer. Obsolescence can be broadly categorized into Planned and Unplanned obsolescence. Planned obsolescence can be considered as a business strategy, in which the obsolescence of a product is built into it from its conception. As Philip Kotler termed it Much so-called planned obsolescence is the working of the competitive and technological forces in a free society-forces that lead to ever-improving goods and services. On the other hand, unplanned obsolescence causes more harm to a burgeoning industry than good. This issue is more prevalent in the electronics industry; the procurem ent life-cycles for electronic components are significantly shorter than the manufacturing and support life-cycle. Therefore, it is highly important to implement and operate an active management of obsolescence to mitigate and avoid extreme costs [1]. One such product that has been plagued by threat of obsolescence is the digital camera. Ever-since the invention of smartphones there has been a huge dip in the digital camera sales, as can be seen from Figure 1. The decreasing price, the exponential rate at which the pixels and the resolution of the smart-phones improved can be termed as few of the factors that cannibalized the digital camera market. Figure 1 Worldwide Sales of Digital Cameras (2011-2016) [2] and Worldwide sale of cellphones on the right (2007-2016) [3] CLUSTERING Humans naturally use clustering to understand the world around them. The ability to group sets of objects based on similarities are fundamental to learning. Researchers have sought to capture these natural learning methods mathematically and this has birthed the clustering research. To help us solve problems at-least approximately as our brain, mathematically precise notation of clustering is important [4]. Clustering is a useful technique to explore natural groupings within multivariate data for a structure of natural groupings, also for feature extraction and summarizing. Clustering is also useful in identifying outliers, forming hypotheses concerning relationships. Clustering can be thought of as partitioning a given space into K groups i.e., à °Ã ââ¬Ëââ¬Å": à °Ã ââ¬Ëâ⬠¹ à ¢Ã¢â¬ ââ¬â¢ {1, à ¢Ã¢â ¬Ã ¦, K}. One method of carrying out this partitioning is to optimize some internal clustering criteria such as the distance between each observation within a c luster etc. While clustering plays an important role in data analysis and serves as a preprocessing step for a multitude of learning task, our primary interest lies in the ability of clusters to gain more information from the data to improve prediction accuracy. As clustering, can be thought of separating classes, it should help in classification task. The aim of clustering is to find useful groups of objects, usefulness being defined by the goals of the data analysis. Most clustering algorithms require us to know the number of clusters beforehand. However, there is no intuitive way of identifying the optimal number of clusters. Identifying optimal clustering is dependent on the methods used for measuring similarities, and the parameters used for partitioning, in general identifying the optimal number of clusters. Determining number of clusters is often an ad hoc decision based on prior knowledge, assumptions, and practical experience is very subjective. This paper performs k-means and k-medoids clustering to gain information from the data structure that could play an important role in predicting obsolescence. It also tries to address the issue of assessing cluster tendency, which is a first and foremost step while carrying out unsupervised machine learning process. Optimization of internal and external clustering criteria will be carried out to identify the optimal number of cluster. Cluster Validation will be carried out to identify the most suitable clustering algorithm. DATA CLEANING Missing value in a dataset is a common occurrence in real world problems. It is important to know how to handle missing data to reduce bias and to produce powerful models. Sometimes ignoring the missing data, biases the answers and potentially leads to incorrect conclusion. Rubin in [7] differentiated between three types of missing values in the dataset: Missing completely at random (MCAR): when cases with missing values can be thought of as a random sample of all the cases; MCAR occurs rarely in practice. Missing at random (MAR): when conditioned on all the data we have, any remaining missing value is completely random; that is, it does not depend on some missing variables. So, missing values can be modelled using the observed data. Then, we can use specialized missing data analysis methods on the available data to correct for the effects of missing values. Missing not at random (MNAR): when data is neither MCAR nor MAR. This is difficult to handle because it will require strong assumptions about the patterns of missing data. While in practice the use of complete case methods which drops the observations containing missing values is quite common, this method has the disadvantage that it is inefficient and potentially leads to bias. Initial approach was to visually explore each individual variable with the help of VIM. However, upon learning the limitations of filling in missing values through exploratory data analysis, this approach was abandoned in favor of multiple imputations. Joint Modelling (JM) and Fully Conditional Specification (FCS) are the two emerging general methods in imputing multivariate data. If multivariate distribution of the missing data is a reasonable assumption, then Joint Modelling which imputes data based on Markov Chain Monte Carlo techniques would be the best method. FCS specifies the multivariate imputation model on a variable-by-variable basis by a set of conditional densities, one for each incomplete variable. Starting from an initial imputation, FCS draws imputations by iterating over the conditional densities. A low number of iterations is often sufficient. FCS is attractive as an alternative to JM in cases where no suitable multivariate distribution can be found [8]. The Multiple imputations approach involves filling in missing values multiple times, creating multiple complete datasets. Because multiple imputations involve creating multiple predictions for each missing value, the analysis of data imputed multiple times take into account the uncertainty in the imputations and yield accurate standard errors. Multiple imputation techniques have been utilized to impute missing values in the dataset, primarily because it preserves the relation in the data and it also preserves uncertainty about these relations. This method is by no means perfect, it has its own complexities. The only complexity was having variables of different types (binary, unordered and continuous), thereby making the application of models, which assumed multivariate normal distribution- theoretically inappropriate. There are several complexities that surface listed in [8]. In order to address this issue It is convenient to specify imputation model separately for each column in th e data. This is called as chained equations wherein the specification occurs at a variable level, which is well understood by the user. The first task is to identify the variables to be included in the imputation process. This generally includes all the variables that will be used in the subsequent analysis irrespective of the presence of missing data, as well as variables that may be predictive of the missing data. There are three specific issues that often come up when selecting variables: (1) creating an imputation model that is more general than the analysis model, (2) imputing variables at the item level vs. the summary level, and (3) imputing variables that reflect raw scores vs. standardized scores. To help make a decision on these aspects, the distribution of the variables may help guide the decision. For example, if the raw scores of a continuous measure are more normally distributed than the corresponding standardized scores then using the raw scores in the imputation model, will likely better meet the assumptions of the linear regressions being used in the imputation process. The following image shows the missing values in the data-frame containing the information regarding digital camera. Figure 2 Missing Variables We can see that Effective Pixels has missing values for all its observations. After cross verifying it with the source website, the web scrapper was rewriting to correctly capture this variable from the website. The date variable was converted from a numeric to a date and this enabled the identification of errors in the observation for USB in the dataset. Two cameras that were released in 1994 1995 were shown to have USB 2.0, after searching online, it was found out that USB 2.0 was released in the year 2005 and USB 1.0 was released in the year 1996. As, most of the cameras before 1997 used PC-serial port a new level was introduced to the USB variable to indicate this. DATA DESCRIPTION The dataset containing the specification of the digital cameras was acquired using rvest -package [5] in R from the url provided in [6]. The structure of the data set is as shown in Appendix A. The data-frame contains 2199 observation and 55 variables. Appendix B contains the descriptive statistics of the quantitative variables in the data-frame. Figure 4 The Distribution of Body-Type in the dataset Observation: Most of the compact, Large SLR and ultracompact cameras are discontinued. Figure 5 Plot showing the status of Digital Cameras from 1994-2017 Observation: Most of the cameras released before 2007 have been discontinued however, we can see that few cameras announced between the period of 1996-2006 are still in production. Fewer new cameras have been announced after the year 2012, this can be evidenced due to the decreasing number of camera sales presented in Figure 5. Figure 6 Distribution of different Cameras (1994-2017) Observation: Between the period of 1996 2012 the digital camera market was dominated by the compact cameras. After 2012, fewer new compact cameras have been announced or are still in production. Same can be said about the fate of ultracompact cameras. In the year 2017, only SLR style mirrorless cameras have been announced, signaling the death of point and shoot cameras. Figure 7 Plot showing the Change in the Total Resolution and Effective Pixels of Digital Camera over the Years Observation: Total resolution has seen an improvement over the years. The presence of outliers can be seen in the top-left corner of the plot. Although the effective pixel is around 10, the total resolution is far higher than any of the cameras announced between the period 1996-2001. These could be the cameras that are still in production as evidenced from Figure 7. ASSESSING CLUSTER TENDENCY A primary issue with unsupervised machine learning is the fact if carried out blindly, clustering methods will divide the data into clusters, because that is what they are supposed to do. Therefore, before choosing a clustering approach, it is important to decide whether the dataset contains meaningful clusters. If the data does contain meaningful clusters, then the number of clusters is also an issue that needs to be looked at. This process is called assessing clustering tendency (feasibility of cluster analysis). To carry out a feasibility study of cluster analysis Hopkins statistic will be used to assess the clustering tendency of the dataset. Hopkins statistic assess the clustering tendency based on the probability that a given data follows a uniform distribution (tests for spatial randomness). If the value of the statistic is close to zero this implies that the data does not follow uniform distribution and thus we can reject the null hypothesis. Hopkins statistic is calculated using the following formula: Where xi is the distance between two neighboring points in a given, dataset and yi represents the distance between two neighboring points of a simulated dataset following uniform distribution. If the value of H is 0.5, this implies that and are close to one another and thus the given data follows a uniform distribution. The next step in the unsupervised learning method is to identify the optimal number of clusters. The Hopkins statistic for the digital camera dataset was found to be 0.00715041. Since Hopkins statistic was quite low, we can conclude that the dataset is highly clusterable. A visual assessment of the clustering tendency was also carried out and the result can be seen in Figure 8. Figure 8 Dissimilarity Matrix of the dataset DETERMINING OPTIMAL NUMBER OF CLUSTERS One simple solution to identify the optimal number of cluster is to perform hierarchical clustering and determine the number of clusters based on the dendogram generated. However, we will utilize the following methods to identify the optimal number of clusters: An optimization criterion such as within sum of squares or Average Silhouette width Comparing evidence against null hypothesis. (Gap Statistic) SUM OF SQUARES The basic idea behind partitioning methods like k-means clustering algorithms, is to define clusters such that the total within cluster sum of squares is minimized. Where Ck is the kth cluster and W(Ck) is the variation within the cluster. Our aim is to minimize the total within cluster sum of squares as it measures the compactness of the clusters. In this approach, we generally perform clustering method, by varying the number of clusters (k). For each k we compute the total within sum of squares. We then plot the total within sum of squares against the k-value, the location of bend or knee in the plot is considered as an appropriate value of the cluster. AVERAGE SILHOUETTE WIDTH Average silhouette is a measure of the quality of clustering, in that it determines the how well an object lies within its cluster. The metric can range from -1 to 1, where higher values are better. Average silhouette method computes the average silhouette of observations for different number of clusters. The optimal number of clusters is the one that maximizes the average silhouette over a range of possible values for different number of clusters [9]. Average silhouette functions similar to within sum of squares method. We carry out the clustering algorithm by varying the number of clusters, then we calculate average silhouette of observation for each cluster. We then plot the average silhouette against different number of clusters. The location with the highest value of average silhouette width is considered as the optimum number of cluster. GAP STATISTIC This method compares the total within sum of squares for different number of cluster with their expected values while assuming that the data follows a distribution with no obvious clustering. The reference dataset is generated using Monte Carlo simulations of the sampling process. For each variable (xi) in the dataset we compute its range [min(xi), max(xj)] and generate n values uniformly from the range min to max. The total within cluster variation for both the observed data and the reference data is computed for different number of clusters. The gap statistic for a given number of cluster is defined as follows: denotes the expectation under a sample of size n from the reference distribution. is defined via bootstrapping and computing the average . The gap statistic measures the deviation of the observed Wk value from its expected value under the null hypothesis. The estimate of the optimal number of clusters will be a value that maximizes Gapn(k). This implies that the clustering structure is far away from the uniform distribution of points. The standard deviation (sdk) of is also computed in order to define the standard error sk as follows: Finally, we choose the smallest value of the number of cluster such that the gap statistic is within one standard deviation of the gap at k+1 Gap(k)à ¢Ã¢â¬ °Ã ¥Gap(k+1) sk+1 The above method and its explanation are borrowed from [10]. DATA PRE-PROCESSING The issue with K-means clustering is that it cannot handle categorical variables. As the K-means algorithm defines a cost function that computes Euclidean distance between two numeric values. However, it is not possible to define such distance between categorical values. Hence, the need to treat categorical data as numeric. While it is not improper to deal with variables in this manner, however categorical variables lose their meaning once they are treated as numeric. To be able to perform clustering efficiently, Gower distance will be used for clustering. The concept of Gower distance is that for each variable a distance metric that works well for that particular type of variable is used. It is scaled between 0 and 1 and then a linear combination of weights is calculated to create the final distance matrix. PARTITIONING METHODS K-MEANS K-means clustering is the simplest and the most commonly used partitioning method for splitting a dataset into a set of k clusters. In this method, we first choose K initial centroids. Each point is then assigned to the closest centroid, and each collection of points is assigned to a centroid in the cluster. The centroid of each cluster is updated based on the additional points assigned to the cluster. We repeat his until the centroids find a steady state. Figure 9 Plot Showing total sum of square and Average Silhouette width for different number of clusters We can see from Figure 9, that the optimal number of clusters suggested by the optimization criteria is 3 clusters using WSS method and 2 clusters using Average Silhouette width method. Considering the dependent variable is factor with two levels, having two clusters does make sense. The disadvantage of optimization criterion to identify the optimal clusters is that, it is sometimes ambiguous. A more sophisticated method is the gap statistic method. Figure 10 Gap Statistic for different number of clusters From Figure 10, we can see that the Gap statistic is high for 2 clusters. Hence, we carry out k-means clustering with 2 clusters on a majority basis. Figure 11 Visualizing K-means Clustering Method The data separates into two relatively distinct clusters, with the red category in the left region, while the region on the right contains the blue category. There is a limited overlap at the interface between the classes. To visualize K-means it is necessary to bring the number of dimensions down to two. The graph produced by fviz_cluster: Factoextra Ver: 1.0 [11] is not a selection of any two dimensions. The plot shows the projection of the entire data onto the first two principle components. These are the dimensions which show the most variation in the data. The 52.8% indicates that the first principle component accounts for 52.8% variation in the data, whereas the second principle component accounts for 23.9% variation in the data. Together both the dimensions account for 76.7% of the variation. The polygon in red and blue represent the cluster means. PARTITIONING AROUND MEDOIDS K means clustering is highly sensitive to outliers, this would affect the assignment of observations to their respective clusters. Partitioning around medoids also known as K-medoids clustering are much more robust compared to k-means. K-medoids is based on the search of medoids among the observation of the dataset. These medoids represent the structure of the data. Much like K-means, after finding the medoids for each of the K- clusters, each observation is assigned to the nearest medoid. The aim is to find K-medoids such that it minimizes the sum of dissimilarities of the observations within the cluster. Figure 12 Plot Showing total sum of square and Average Silhouette width for different number of clusters We can see from Figure 12, that the optimal number of clusters suggested by the optimization criteria is 3 clusters using WSS method and 2 clusters using Average Silhouette width method. Considering the dependent variable is factor with two levels, having two clusters does make sense. The disadvantage of optimization criterion to identify the optimal clusters is that, it is sometimes ambiguous. A more sophisticated method is the gap statistic method. Figure 13 Gap Statistic for different number of clusters From Figure 13, we can see that the Gap statistic is high for 2 clusters. Hence, we carry out partitioning around medoids clustering with 2 clusters on a majority basis. Figure 14 Plot visualizing PAM clustering method The data separates into two relatively distinct clusters, with the red category in the lower region, while the upper region contains the blue category. There is a limited overlap at the interface between the classes. fviz_cluster: Factoextra Ver: 1.0 [11] transforms the initial set of variables into a new set of variables through principal component analysis. This dimensionality reduction algorithm operates on the 72 variables and outputs the two new variables that represent the projection of the original dataset. CLUSTER VALIDATION The next step in cluster analysis is to find the goodness of fit and to avoid finding patterns in noise and to compare clustering algorithms, cluster validation is carried out. The following cluster validation measures to compare K-means and PAM clustering will be used: Connectivity: Indicates the extent to which the observations are placed in the same cluster as their nearest neighbors in the data space. It has a value ranging from 0 to à ¢Ãâ Ã
¾ and should be minimized Dunn: It is the ratio of shortest distance between two clusters to the largest intra-cluster distance. It has a value ranging from 0 to à ¢Ãâ Ã
¾ and should be maximized. Average Silhouette width The results of internal validation measures are presented in the table below. K-means for two cluster has performed better for each statistic. Figure 15 Plot Comparing Connectivity and Dunn Index for K-means and PAM for different number of clusters à à à Figure 16 Plot Comparing Average Silhouette width of K-means and PAM Clustering Algorithm Validation Measures Number of Clusters 2 3 4 5 6 kmeans Connectivity 139.9575 292.5563 406.5429 514.3913 605.5373 Dunn 0.0661 0.0246 0.0223 0.0244 0.0291 Silhouette 0.4369 0.3174 0.2814 0.2679 0.2447 pam Connectivity 156.1004 333.754 474.4298 520.3913 635.3687 Dunn 0.0275 0.0397 0.022 0.028 0.0246 Silhouette 0.4271 0.3035 0.2757 0.2661 0.2325 Table 1 Presenting the values of different validation measures for K-means and PAM Validation Measures Score Method Clusters Connectivity 139.9575 kmeans 2 Dunn 0.0661 kmeans 2 Silhouette 0.4369 kmeans 2 Table 2 Optimal Scores for the Validation Measures CONCLUSION In this research work, partitioning methods like K-means and Partitioning around medoids were developed. The performances of these two approaches have been observed on the basis of their Connectivity, Dunn index and Average Silhouette width. The results indicate that K-means clustering algorithm with K = 2 performs better than partitioning around medoids with two clusters. The findings of this paper will be very useful to predict obsolescence with higher accuracy. FUTURE WORK Advanced clustering algorithms such as Model based clustering and Density based clustering can be carried out to find the multivariate data structure as most of the variables are categorical. [1] Bjoern Bartels, Ulrich Ermel, Peter Sandborn and Michael G. Pecht (2012). Strategies to the Prediction, Mitigation and Management of Product Obsolescence. [2] Source Figure 1: https://www.statista.com/statistics/269927/sales-of-analog-and-digital-cameras-worldwide-since-2002/ [3] Source, Figure 1: https://www.statista.com/statistics/263437/global-smartphone-sales-to-end-users-since-2007/ [4] S. Still, and W. Bialek, How many Clusters? An Information Theoretic Perspective, Neural Computation, 2004. [5] Wickham, Hadley, rvest: Easily Harvest (Scrape) Web Pages. https://cran.r-project.org/web/packages/rvest/rvest.pdf, Ver. 0.3.2 [6] https://www.dpreview.com [7] Rubin, D.B., Inference and missing data. Biometrika, 1976. [8] Multivariate Imputation by Chained Equations Stef van Buuren, Karin Groothuis . [9] Learning the k in k-means Greg Hamerly, Charles Elkan [10] Robert Tibshirani, Guenther Walther and Trevor Hast
Sunday, January 19, 2020
Assess the usefulness of social action theories in the study of society Essay
Social action theories are known as micro theories which take a bottom-up approach to studying society; they look at how individuals within society interact with each other. There are many forms of social action theories, the main ones being symbolic interactionism, phenomenology and ethnomethodology. They are all based on the work of Max Weber, a sociologist, who acknowledged that structural factors can shape our behaviour but individuals do have reasons for their actions. He used this to explain why people behave in the way in which they do within society. Weber saw four types of actions which are commonly committed within society; rational, this includes logical plans which are used to achieve goals, traditional-customary behaviour, this is behaviour which is traditional and has always been done; he also saw affectual actions, this includes an emotion associated with an action and value-rational actions, this is behaviour which is seen as logical by an individual. Weberââ¬â¢s discovery of these actions can therefore be seen as useful in the study of society. Weber discovered these actions by using his concept of verstehan, a deeper understanding. However, some sociologists have criticised him as they argue that verstehan cannot be accomplished as it is not possible to see thing in the way that others see them, leaving sociologists to question whether Weberââ¬â¢s social action theory is useful in the study of society. Social action theories have also been referred to as interactionism as they aim to explain day-to-day interactions between individuals within society. G. H Mead came up with the idea of interactionism and argued that the self is ââ¬Ëa social construction arising out of social experienceââ¬â¢. This is because, according to Mead, social situations are what influence the way in we act and behave. He claims that we develop a sense of self as a child and this allows us to see ourselves in the way in which other people see us; we act and behave in certain ways depending on the circumstances which we are in. Mead also claimed that we have a number of different selves which we turn into when we are in certain situations; i. e. we may have one self for the work place and another self for home life. Mead concluded that society is like a stage, in which we are all ââ¬Ëactorsââ¬â¢. Meadââ¬â¢s theory if interactionism is useful in the study of society as it explains why people behave in different ways in certain situations. Mead argues that the social context of a situation is what influences our behaviour, humans use symbols, in the form of language and facial expressions, to communicate, he also argued that humans and animals differ as reasons behind humansââ¬â¢ actions are thought through and not instinctive, unlike those of animalsââ¬â¢. However, it has been argued that not all action is meaningful, as Weberââ¬â¢s category of traditional action suggests that much action is performed unconsciously and may have little meaning. Therefore, meadââ¬â¢s idea of interactionism cannot be seen as an appropriate theory to use when studying society. Blumer, a sociologist, who elaborated on Meadââ¬â¢s concept of the self ââ¬â ââ¬ËIââ¬â¢ and ââ¬Ëmeââ¬â¢ ââ¬â stated that there were three principles about actions and behaviours within social situations. He argued that our actions are the result of situations and events and they have reasons. The reasons behind our actions are negotiable and changeable, so theyââ¬â¢re not fixed. Our interpretation of a situation is what gives it meaning. Blumerââ¬â¢s three principles can therefore be used in the study of society. However, it has been argued that his principles cannot explain the consistent patterns which we see in peopleââ¬â¢s behaviours. This therefore leaves many sociologists to question whether Blumerââ¬â¢s principles can be used to study society. Labelling theory has also been used to apply the interactionist theory to society; the theory, like Mead, emphasises the importance of symbols and situations in which they are used. The main interactionist concepts are the definition of the situation ââ¬â if we believe in something then it could affect the way in which we behave. The looking glass ââ¬âself ââ¬â this was created by Cooley who argues that we see ourselves in a way in which we think others see us. These concepts have been useful in explaining why people act in certain ways in certain situations; therefore, the labelling theory is effective in the study of society. Overall, in conclusion, there are many different social action theories which can be used in the study of society, however, not all of them can be applied to all individuals.
Saturday, January 11, 2020
Network Server Administration
Course number CIS 332, Network Server Administration, lists as its main topics: installing and configuring servers, network protocols, resource and end user management, security, Active Directory, and the variety of server roles which can be implemented. My experience and certification as a Microsoft Certified System Administrator (MCSA) as well as a Microsoft Certified System Engineer (MCSE) demonstrates that I have a thorough grounding in both the theory and practice of the topics covered in this course and should receive credit for it. Installing and configuring servers was the subject of Installing, Configuring and Administering Microsoft Windows 2000 Server, which I took in 2001 in preparation for my initial Microsoft Certified Professional certification. This exam covered such topics as installing Microsoft Windows 2000 Server using both an attended installation and an unattended installation; server upgrades from Windows NT (the previous version) and troubleshooting and repairing failed installations. This exam also covered installing and configuring hardware devices and user management. Network protocols were discussed during the training for the exam Implementing and Administering a Microsoft Windows 2000 Network Infrastructure, which I also took in 2001. This exam covered installing, configuring, troubleshooting and administering such protocols as DNS and DHCP, TCP/IP, NWLink, and IPSec. The training covered such aspects of network protocols as remote access policies and network routing. Security was one of the topics of this exam, as well. Network security using IPSec and encryption and authentication protocols was discussed along with the network implementation details. Resource and end user management was one of the main topics of the Managing and Maintaining a Windows Server 2003 Environment exam, which also updated my knowledge of security, networking and utilities. The exam covered such topics as user creation and modification, user and group management, Terminal Services management and implementing security and software update services. Security was covered in a number of exams, including Implementing and Administering a Microsoft Windows 2000 Network Infrastructure, Installing, Configuring and Administering Microsoft Windows 2000 Server and Designing Security forà a Windows 2000 Network. All aspects of network security were covered in the various training sessions for these exams, including topics such as analysis of network security requirements in relation to organizational realities and requirements, design and implementation of such specifics as authentication policies, public-key infrastructures and encryption techniques, physical security, and design and implementation of security audit and assurance strategies. Also included were security considerations for all auxiliary services, such as DNS, Terminal Services, SNMP, Remote Installation Services and others. Implementation of Active Directory and knowledge of varied server roles was provided by the exam Designing a Microsoft Windows 2000 Directory Services Infrastructure. The training for this exam encompassed the design and implementation of an Active Directory forest and domain structure as well as planning a DNS strategy for client and server naming. This training also included design and implementation of a number of different server types, such as file and print servers, databases, proxy servers, Web servers, desktop management servers, applications servers and dial-in management servers. Further knowledge of Active Directory and auxiliary services was provided in the training for Implementing and Administering a Microsoft Windows 2000 Directory Services Infrastructure. This training included such topics as installing, configuring and troubleshooting Active Directory and DNS, implementing Change and Configuration Management, and managing all the components of Active Directory, including moving, publishing and locating Active Directory Objects, controlling access, delegating administrative privileges for objects, performing backup and restore and maintaining security for the Active Directory server via Group Policy and the Security Configuration and Analysis tool. The topics covered in CIS 332, Network Server Administration, have been completely encompassed by my previous experience, training and certification with Microsoft Windows Server 2000, as well as updated knowledge gained byà training for Microsoft Windows Server 2003. I have been constantly increasing my skills and knowledge in this area for the past six years, using both training and work experience to gain certifications which prove that I have a complete grasp of all aspects of the subject matter included in this course. Installing and configuring servers and network protocols, troubleshooting failed installations or configurations, resource and end user management, security design and management, design and implementation of Active Directory services and implementing and administering a wide variety of network server roles are all major aspects of my training and certification experience. I feel I am fully qualified for the information covered in CIS 332, and should be granted credit for this course.
Subscribe to:
Posts (Atom)