This is the unclassified or linear scale for the data. Unclassified maps show data in a range of values from the minimum to the maximum value. For this data set, the minimum value is 0 and goes to 13%. This shows where the majority of values lie; for this, it is in the lower values. Although when looking at an unclassified map, it can be difficult to differentiate the middle values from eachother making the map more difficult to interpret. In observable, unclassified is called linear so we have to use the entire extent or range of values for the domain, while the range is depicted going from while to blue.
To create this color scale, you have to input two colors, the most effective being a scale from white to a concentrated color into the range. For colors, you can either put basic color names that Observable has downloaded into the data or you can put a color code found from the internet, which is common for all coding.
Quantile classification is shown on a scale where the domain is the value being classified: the invasive percent. Quantile classification divides the total number of observations by the number of bins to set an equal number of observations in each bin. The range is then classified based on the observations. For this data set, we have 64 observations or counties so each bin has around 21 observations. The range of data in each bin depicts the values that the observations contain, so the scale won't necessarily be even in terms of value.
When using a certain number of bins, it is important to make sure each bin has a different color, and it is generally more effective to have lighter colors show lower values and more saturated or darker colors show higher values. For Obsevable, the colors are listed in the range where for this data, white shows the first bin, light blue shows the second bin, and steelblue shows the third bin. Where white has the lowest percent, lightblue has the midrange percent, and steelblue has the largest percent of values.
Jenks natural breaks classification classifies the data based on the natural breaks within the data. This method minimises differences in values within the same class and maximises the difference between classes. For a right-skewed distribution of data, like the data I am using, this means the majority of observations reside within the first class.
for the data input in observable, which follows the same general layout for each type of classification.
Jenks- the type of classification that observable recognizes, and the general name
= d3- number of bins
.scaleThreshold- the type of scale for each classification that observable recognizes
.domain(naturalbreaks) - where the values lie on the scale and how they are broken up into classes
.range () -what colors the bins are assigned in order
To best show my data, the percentage of invasive plants per Colorado county I set the number of bins to be 7, while the range showed the entire data set from 0 to 13%. I found that with an increased number of bins, there were breaks within the data or outliers where the percentage was higher than the other data points. Like for my data, some counties had 9% cover of invasive plants, then skipped to 11%. So I decided to make the range of values in each bin larger since the outliers of an increased two percent really aren't that significant. Since my data is right-skewed, I wanted to make sure the data showed a decrease in percentage without non-significant outliers existing. No matter how many bins I created, the data still showed a linear decrease in percentage, with the large majority of values being in the smallest bin.
Determining the number of classes is important and varies based on data. Having too few classes might not necessarily get the point across in what you're trying to show to viewers, and having too many classes can also have the same effect, but make it a little more difficult to interpret. It's good to test this out to see how the data is aggregated when increasing and decreasing class numbers to make sure the intended message of the data is most effectively being interpreted.
Equal interval classification uses the data values' range as a domain from the minimum to the maximum value and splits these values into the number of classes given. In this case, there are three classes. So the first one would range from 0 -4.3%, then 4.4-8.6%, then 8.7-13%. This makes sure that the range of percentages for each class is equal. This does not take into account the number of observations or show where most observations lie.
In observable Threshold is a manual classification where you can input the maximum value for each set number of bins in the domain. Since the majority of my observations had a percentage of invasive plants between 0-2% I wanted to show that as one of my classes. Next, I took the remaining percent and divided it by around 2 to show the highest observations that have a larger percent than the middle value.
Showing how the data is distributed in a horizontal bar chart for each of the classification methods is very effective, especially when the data is being compared at the top. You can see that when using a quantile scale for my data, making the number of observations equal shows that the range for each class exponentially increases. You can see how for quantize or equal interval, the range of values is perfectly divided into each class, which makes a large number of data points fall into the first class. For Jenks, when comparing it to the data, the observations are fairly distributed throughout each class. And how manually inputting the range of values for each class, threshold, is a good in-between of all the different types of classification.