含羞草能治什么病| 521是什么意思| 小刺猬吃什么东西| 为什么白醋把纹身洗掉了| 京酱肉丝用什么酱| 仕途是什么意思| 男生喜欢什么礼物| 鲁迅的原名叫什么| 肌红蛋白低说明什么| 皮肤黑穿什么颜色| 冷面是什么面| 鳞状上皮内高度病变是什么意思| 肠胃不好适合喝什么茶| 硬度不够吃什么药调理| 云是由什么组成的| 母亲节送妈妈什么| 毛孔粗大用什么洗面奶好| 房室传导阻滞是什么意思| yet什么意思| 什么叫造口| 飞机下降时耳朵疼是什么原因| 1974年属虎是什么命| 唐僧念的紧箍咒是什么| 7月出生的是什么星座| 森林里有什么| 海带和什么菜搭配好吃| 指征是什么意思| herry是什么意思| 什么水果不能吃| 纤支镜检查是用来查什么的| 什么的图案| 吃什么降血脂和胆固醇| 什么什么不舍| 为什么会有湿疹| 人吸了甲醛有什么症状| 10月16日什么星座| 记忆力下降是什么原因引起的| 吃避孕药有什么危害| 督察是什么意思| 喜欢吃酸的是什么原因| 孕晚期吃什么好| 空气棉是什么面料| 什么的绽放| 根茎叶属于什么器官| 屁股疼什么原因| kim是什么意思| 十滴水泡脚有什么好处| 美国为什么帮以色列| 家里为什么会进蝙蝠| 五蕴皆空是什么意思| 腹胀屁多是什么原因| 嘴巴发麻是什么原因| 娃儿发烧用什么方法退烧快| 高足是什么意思| 血糖高的人早餐吃什么最好| 胃不舒服想吐吃什么药| 子母环是什么形状图片| 斜杠青年什么意思| 8月6号什么星座| 手上的月牙代表什么意思| 男人很man是什么意思| betty是什么意思| 舌苔发黑是什么病的前兆| 肠胃炎吃什么抗生素| 健身前吃什么比较好| s 是什么意思| 医保乙类是什么意思| 为什么一喝牛奶就拉肚子| 虾不能和什么同吃| dove什么意思| 趾高气昂是什么意思| 喜欢是什么感觉| 男人阳气不足有什么症状| 什么花秋天开| 血常规crp是什么意思| 蛋白粉适合什么人群吃| 蘑菇和什么不能一起吃| 1946年属什么生肖属相| 戏谑什么意思| 红色代表什么| 梦见买馒头是什么意思| 侯字五行属什么| 胎菊泡水喝有什么功效| 125是什么意思| 十二生肖它第一是什么生肖| 时隔是什么意思| 中产阶级的标准是什么| 1978年是什么年| 嗓子有异物感吃什么药| 87年是什么年| 痛风打什么针| 彩色多普勒超声常规检查是什么| 背疼挂什么科室最好| 什么叫高危行为| 怀孕吸烟对胎儿有什么影响| 粘纤是什么面料优缺点| 承你吉言是什么意思| 调制乳粉是什么意思| 慢性结肠炎用什么药| 不速之客的速是什么意思| 梦见摘杏子是什么意思| 空调不出水是什么原因| 为什么会肛裂| 鱼油什么牌子好| 八一年属什么生肖| 直言不讳是什么意思| 人为什么要抽烟| 小孩便秘吃什么食物好| 转学需要什么手续| 前列腺肥大吃什么药效果最好| 什么是紫外线| 生产批号是什么意思| 萝卜喝醉了会变成什么| 见路不走是什么意思| 肠炎可以吃什么水果| 乳腺囊肿有什么症状| 幽门螺旋杆菌感染有什么症状| 好奇害死猫是什么意思| 合成碳硅石是什么| 流注是什么意思| 什么是蚂蚁上树| 出血热是什么病| 今年是农历的什么年| 小舅子是什么意思| 免疫球蛋白e高说明什么| 梦见打死黄鼠狼是什么意思| 抗心磷脂抗体是什么| 孔子是什么学派的创始人| 离殇是什么意思| 梦见大便是什么意思| honey什么意思| 梦见红色的蛇是什么意思| 可颂是什么意思| 电销是什么工作| 尿酸高不能吃什么蔬菜| 下午1点到3点是什么时辰| 痧是什么| 蚝油可以用什么代替| 杂交金毛犬长什么样子| 什么馅饺子好吃| 消炎痛又叫什么| 前列腺增大伴钙化灶是什么意思| 重阳节送老人什么礼物| 脆豆腐是什么做的| 竹叶青属于什么茶| 脚气用什么药膏效果好| hm平方是什么单位| 7.6什么星座| iphone5什么时候出的| 老日念什么| 梦见自己抬棺材是什么意思| 贾字五行属什么| 益生菌适合什么人群吃| 墨龟为什么只能养一只| b超和彩超有什么区别| 戒指戴左手中指是什么意思| 巴基斯坦用什么语言| 胃疼喝什么可以缓解| 枕戈待旦什么意思| 欢字五行属什么| 百雀羚适合什么年龄段| 不止是什么意思| 广东有什么特色美食| 果胶是什么| 梦见死蛇是什么预兆| abby是什么意思| gg了是什么意思| 狮子座和什么星座不合| 情有独钟是什么意思| 吃什么代谢快| n2是什么意思| 826是什么星座| 市宣传部长是什么级别| 什么叫屌丝| 银耳不能和什么一起吃| 老鹰代表什么生肖| 加拿大现在是什么时间| 嗓子有异物感堵得慌吃什么药| 奔现是什么意思| c5是什么意思| 鹰头皮带是什么牌子| 为什么大医院不用宫腔镜人流| 月什么意思| 社保缴费基数和工资有什么关系| 肾病应该吃什么| 七月十六是什么星座| 替代品是什么意思| 拖是什么意思| 裂纹舌是什么原因引起的| 气管炎用什么药| 怀孕腿抽筋是因为什么原因引起的| 什么是开悟| 吃什么补维生素| 九七年属什么| 什么情况要打破伤风针| 生姜放肚脐眼有什么功效| 你好是什么意思| 肚脐下方疼是什么原因| 3月17日是什么星座| 浅棕色是什么颜色| 喉咙一直有痰是什么原因| 米老鼠叫什么名字| 个个想出头是什么生肖| 五脏六腑指什么| 乾字五行属什么| 阑尾有什么作用| 为什么膝盖弯曲就疼痛| 血脂稠喝什么茶效果好| 喜爱的反义词是什么| 低血压吃什么药效果好| 气性大是什么意思| 梦到黄鳝是什么意思| swissmade是什么意思| 子宫肌瘤吃什么药| 质感是什么意思| 什么扑鼻| 肝肾功能检查挂什么科| 斑秃吃什么药效果好| 1972年属什么| 东倒西歪的动物是什么生肖| 伤骨头了吃什么好得快| 探望病人买什么水果| 突然胃疼是什么原因| 什么叫失眠| 腐竹是什么做的| 直肠炎吃什么药效果好| 查宝宝五行八字缺什么| 瑞舒伐他汀钙片治什么病| 装什么病能容易开病假| bpo是什么意思啊| 尿结石挂什么科| 卉是什么意思| tomboy什么意思| 肺部条索灶是什么意思| ova什么意思| 肿瘤前期有什么症状| 长针眼是什么意思| 附件炎吃什么药效果好| 鬼代表什么数字| 什么是气短| 吃什么食物养肝护肝| 久坐腰疼是什么原因| 6月18日是什么节| 女人熬夜吃什么抗衰老| 农历八月初一是什么星座| 乳腺增生吃什么药最好| 周莹是什么电视剧| 绿茶什么时候喝最好| 什么什么致志| 肠胃炎吃什么药好得快| 他达拉非片是什么药| 一什么天安门| 九加虎念什么| 孕妇感染弓形虫有什么症状| 菠萝有什么功效和作用| 兆以上的计数单位是什么| 吃什么可以抗衰老| 什么手机拍照效果最好| 蟑螂喜欢吃什么| 一进大门看见什么最好| 儿茶是什么中药| 什么是隐血| 娃娃鱼属于什么类动物| 成吉思汗姓什么| 百度Jump to content

助力全域旅游示范省建设 省公路管理局引进一批...

From Wikipedia, the free encyclopedia
百度 由于京津冀及周边地区以重化工为主的产业结构、以公路为主的交通运输结构特点,主要污染物排放强度仍处在高位。

In statistics, a categorical variable (also called qualitative variable) is a variable that can take on one of a limited, and usually fixed, number of possible values, assigning each individual or other unit of observation to a particular group or nominal category on the basis of some qualitative property.[1] In computer science and some branches of mathematics, categorical variables are referred to as enumerations or enumerated types. Commonly (though not in this article), each of the possible values of a categorical variable is referred to as a level. The probability distribution associated with a random categorical variable is called a categorical distribution.

Categorical data is the statistical data type consisting of categorical variables or of data that has been converted into that form, for example as grouped data. More specifically, categorical data may derive from observations made of qualitative data that are summarised as counts or cross tabulations, or from observations of quantitative data grouped within given intervals. Often, purely categorical data are summarised in the form of a contingency table. However, particularly when considering data analysis, it is common to use the term "categorical data" to apply to data sets that, while containing some categorical variables, may also contain non-categorical variables. Ordinal variables have a meaningful ordering, while nominal variables have no meaningful ordering.

A categorical variable that can take on exactly two values is termed a binary variable or a dichotomous variable; an important special case is the Bernoulli variable. Categorical variables with more than two possible values are called polytomous variables; categorical variables are often assumed to be polytomous unless otherwise specified. Discretization is treating continuous data as if it were categorical. Dichotomization is treating continuous data or polytomous variables as if they were binary variables. Regression analysis often treats category membership with one or more quantitative dummy variables.

Examples of categorical variables

[edit]

Examples of values that might be represented in a categorical variable:

  • Demographic information of a population: gender, disease status.
  • The blood type of a person: A, B, AB or O.
  • The political party that a voter might vote for, e.?g. Green Party, Christian Democrat, Social Democrat, etc.
  • The type of a rock: igneous, sedimentary or metamorphic.
  • The identity of a particular word (e.g., in a language model): One of V possible choices, for a vocabulary of size V.

Notation

[edit]

For ease in statistical processing, categorical variables may be assigned numeric indices, e.g. 1 through K for a K-way categorical variable (i.e. a variable that can express exactly K possible values). In general, however, the numbers are arbitrary, and have no significance beyond simply providing a convenient label for a particular value. In other words, the values in a categorical variable exist on a nominal scale: they each represent a logically separate concept, cannot necessarily be meaningfully ordered, and cannot be otherwise manipulated as numbers could be. Instead, valid operations are equivalence, set membership, and other set-related operations.

As a result, the central tendency of a set of categorical variables is given by its mode; neither the mean nor the median can be defined. As an example, given a set of people, we can consider the set of categorical variables corresponding to their last names. We can consider operations such as equivalence (whether two people have the same last name), set membership (whether a person has a name in a given list), counting (how many people have a given last name), or finding the mode (which name occurs most often). However, we cannot meaningfully compute the "sum" of Smith + Johnson, or ask whether Smith is "less than" or "greater than" Johnson. As a result, we cannot meaningfully ask what the "average name" (the mean) or the "middle-most name" (the median) is in a set of names.

This ignores the concept of alphabetical order, which is a property that is not inherent in the names themselves, but in the way we construct the labels. For example, if we write the names in Cyrillic and consider the Cyrillic ordering of letters, we might get a different result of evaluating "Smith < Johnson" than if we write the names in the standard Latin alphabet; and if we write the names in Chinese characters, we cannot meaningfully evaluate "Smith < Johnson" at all, because no consistent ordering is defined for such characters. However, if we do consider the names as written, e.g., in the Latin alphabet, and define an ordering corresponding to standard alphabetical order, then we have effectively converted them into ordinal variables defined on an ordinal scale.

Number of possible values

[edit]

Categorical random variables are normally described statistically by a categorical distribution, which allows an arbitrary K-way categorical variable to be expressed with separate probabilities specified for each of the K possible outcomes. Such multiple-category categorical variables are often analyzed using a multinomial distribution, which counts the frequency of each possible combination of numbers of occurrences of the various categories. Regression analysis on categorical outcomes is accomplished through multinomial logistic regression, multinomial probit or a related type of discrete choice model.

Categorical variables that have only two possible outcomes (e.g., "yes" vs. "no" or "success" vs. "failure") are known as binary variables (or Bernoulli variables). Because of their importance, these variables are often considered a separate category, with a separate distribution (the Bernoulli distribution) and separate regression models (logistic regression, probit regression, etc.). As a result, the term "categorical variable" is often reserved for cases with 3 or more outcomes, sometimes termed a multi-way variable in opposition to a binary variable.

It is also possible to consider categorical variables where the number of categories is not fixed in advance. As an example, for a categorical variable describing a particular word, we might not know in advance the size of the vocabulary, and we would like to allow for the possibility of encountering words that we have not already seen. Standard statistical models, such as those involving the categorical distribution and multinomial logistic regression, assume that the number of categories is known in advance, and changing the number of categories on the fly is tricky. In such cases, more advanced techniques must be used. An example is the Dirichlet process, which falls in the realm of nonparametric statistics. In such a case, it is logically assumed that an infinite number of categories exist, but at any one time most of them (in fact, all but a finite number) have never been seen. All formulas are phrased in terms of the number of categories actually seen so far rather than the (infinite) total number of potential categories in existence, and methods are created for incremental updating of statistical distributions, including adding "new" categories.

Categorical variables and regression

[edit]

Categorical variables represent a qualitative method of scoring data (i.e. represents categories or group membership). These can be included as independent variables in a regression analysis or as dependent variables in logistic regression or probit regression, but must be converted to quantitative data in order to be able to analyze the data. One does so through the use of coding systems. Analyses are conducted such that only g -1 (g being the number of groups) are coded. This minimizes redundancy while still representing the complete data set as no additional information would be gained from coding the total g groups: for example, when coding gender (where g = 2: male and female), if we only code females everyone left over would necessarily be males. In general, the group that one does not code for is the group of least interest.[2]

There are three main coding systems typically used in the analysis of categorical variables in regression: dummy coding, effects coding, and contrast coding. The regression equation takes the form of Y = bX + a, where b is the slope and gives the weight empirically assigned to an explanator, X is the explanatory variable, and a is the Y-intercept, and these values take on different meanings based on the coding system used. The choice of coding system does not affect the F or R2 statistics. However, one chooses a coding system based on the comparison of interest since the interpretation of b values will vary.[2]

Dummy coding

[edit]

Dummy coding is used when there is a control or comparison group in mind. One is therefore analyzing the data of one group in relation to the comparison group: a represents the mean of the control group and b is the difference between the mean of the experimental group and the mean of the control group. It is suggested that three criteria be met for specifying a suitable control group: the group should be a well-established group (e.g. should not be an "other" category), there should be a logical reason for selecting this group as a comparison (e.g. the group is anticipated to score highest on the dependent variable), and finally, the group's sample size should be substantive and not small compared to the other groups.[3]

In dummy coding, the reference group is assigned a value of 0 for each code variable, the group of interest for comparison to the reference group is assigned a value of 1 for its specified code variable, while all other groups are assigned 0 for that particular code variable.[2]

The b values should be interpreted such that the experimental group is being compared against the control group. Therefore, yielding a negative b value would entail the experimental group have scored less than the control group on the dependent variable. To illustrate this, suppose that we are measuring optimism among several nationalities and we have decided that French people would serve as a useful control. If we are comparing them against Italians, and we observe a negative b value, this would suggest Italians obtain lower optimism scores on average.

The following table is an example of dummy coding with French as the control group and C1, C2, and C3 respectively being the codes for Italian, German, and Other (neither French nor Italian nor German):

Nationality C1 C2 C3
French 0 0 0
Italian 1 0 0
German 0 1 0
Other 0 0 1

Effects coding

[edit]

In the effects coding system, data are analyzed through comparing one group to all other groups. Unlike dummy coding, there is no control group. Rather, the comparison is being made at the mean of all groups combined (a is now the grand mean). Therefore, one is not looking for data in relation to another group but rather, one is seeking data in relation to the grand mean.[2]

Effects coding can either be weighted or unweighted. Weighted effects coding is simply calculating a weighted grand mean, thus taking into account the sample size in each variable. This is most appropriate in situations where the sample is representative of the population in question. Unweighted effects coding is most appropriate in situations where differences in sample size are the result of incidental factors. The interpretation of b is different for each: in unweighted effects coding b is the difference between the mean of the experimental group and the grand mean, whereas in the weighted situation it is the mean of the experimental group minus the weighted grand mean.[2]

In effects coding, we code the group of interest with a 1, just as we would for dummy coding. The principal difference is that we code ?1 for the group we are least interested in. Since we continue to use a g - 1 coding scheme, it is in fact the ?1 coded group that will not produce data, hence the fact that we are least interested in that group. A code of 0 is assigned to all other groups.

The b values should be interpreted such that the experimental group is being compared against the mean of all groups combined (or weighted grand mean in the case of weighted effects coding). Therefore, yielding a negative b value would entail the coded group as having scored less than the mean of all groups on the dependent variable. Using our previous example of optimism scores among nationalities, if the group of interest is Italians, observing a negative b value suggest they obtain a lower optimism score.

The following table is an example of effects coding with Other as the group of least interest.

Nationality C1 C2 C3
French 0 0 1
Italian 1 0 0
German 0 1 0
Other ?1 ?1 ?1

Contrast coding

[edit]

The contrast coding system allows a researcher to directly ask specific questions. Rather than having the coding system dictate the comparison being made (i.e., against a control group as in dummy coding, or against all groups as in effects coding) one can design a unique comparison catering to one's specific research question. This tailored hypothesis is generally based on previous theory and/or research. The hypotheses proposed are generally as follows: first, there is the central hypothesis which postulates a large difference between two sets of groups; the second hypothesis suggests that within each set, the differences among the groups are small. Through its a priori focused hypotheses, contrast coding may yield an increase in power of the statistical test when compared with the less directed previous coding systems.[2]

Certain differences emerge when we compare our a priori coefficients between ANOVA and regression. Unlike when used in ANOVA, where it is at the researcher's discretion whether they choose coefficient values that are either orthogonal or non-orthogonal, in regression, it is essential that the coefficient values assigned in contrast coding be orthogonal. Furthermore, in regression, coefficient values must be either in fractional or decimal form. They cannot take on interval values.

The construction of contrast codes is restricted by three rules:

  1. The sum of the contrast coefficients per each code variable must equal zero.
  2. The difference between the sum of the positive coefficients and the sum of the negative coefficients should equal 1.
  3. Coded variables should be orthogonal.[2]

Violating rule 2 produces accurate R2 and F values, indicating that we would reach the same conclusions about whether or not there is a significant difference; however, we can no longer interpret the b values as a mean difference.

To illustrate the construction of contrast codes consider the following table. Coefficients were chosen to illustrate our a priori hypotheses: Hypothesis 1: French and Italian persons will score higher on optimism than Germans (French = +0.33, Italian = +0.33, German = ?0.66). This is illustrated through assigning the same coefficient to the French and Italian categories and a different one to the Germans. The signs assigned indicate the direction of the relationship (hence giving Germans a negative sign is indicative of their lower hypothesized optimism scores). Hypothesis 2: French and Italians are expected to differ on their optimism scores (French = +0.50, Italian = ?0.50, German = 0). Here, assigning a zero value to Germans demonstrates their non-inclusion in the analysis of this hypothesis. Again, the signs assigned are indicative of the proposed relationship.

Nationality C1 C2
French +0.33 +0.50
Italian +0.33 ?0.50
German ?0.66 0

Nonsense coding

[edit]

Nonsense coding occurs when one uses arbitrary values in place of the designated "0"s "1"s and "-1"s seen in the previous coding systems. Although it produces correct mean values for the variables, the use of nonsense coding is not recommended as it will lead to uninterpretable statistical results.[2]

Embeddings

[edit]

Embeddings are codings of categorical values into low-dimensional real-valued (sometimes complex-valued) vector spaces, usually in such a way that ‘similar’ values are assigned ‘similar’ vectors, or with respect to some other kind of criterion making the vectors useful for the respective application. A common special case are word embeddings, where the possible values of the categorical variable are the words in a language and words with similar meanings are to be assigned similar vectors.

Interactions

[edit]

An interaction may arise when considering the relationship among three or more variables, and describes a situation in which the simultaneous influence of two variables on a third is not additive. Interactions may arise with categorical variables in two ways: either categorical by categorical variable interactions, or categorical by continuous variable interactions.

Categorical by categorical variable interactions

[edit]

This type of interaction arises when we have two categorical variables. In order to probe this type of interaction, one would code using the system that addresses the researcher's hypothesis most appropriately. The product of the codes yields the interaction. One may then calculate the b value and determine whether the interaction is significant.[2]

Categorical by continuous variable interactions

[edit]

Simple slopes analysis is a common post hoc test used in regression which is similar to the simple effects analysis in ANOVA, used to analyze interactions. In this test, we are examining the simple slopes of one independent variable at specific values of the other independent variable. Such a test is not limited to use with continuous variables, but may also be employed when the independent variable is categorical. We cannot simply choose values to probe the interaction as we would in the continuous variable case because of the nominal nature of the data (i.e., in the continuous case, one could analyze the data at high, moderate, and low levels assigning 1 standard deviation above the mean, at the mean, and at one standard deviation below the mean respectively). In our categorical case we would use a simple regression equation for each group to investigate the simple slopes. It is common practice to standardize or center variables to make the data more interpretable in simple slopes analysis; however, categorical variables should never be standardized or centered. This test can be used with all coding systems.[2]

See also

[edit]

References

[edit]
  1. ^ Yates, Daniel S.; Moore, David S.; Starnes, Daren S. (2003). The Practice of Statistics (2nd ed.). New York: Freeman. ISBN 978-0-7167-4773-4. Archived from the original on 2025-08-07. Retrieved 2025-08-07.
  2. ^ a b c d e f g h i j Cohen, J.; Cohen, P.; West, S. G.; Aiken, L. S. (2003). Applied multiple regression/correlation analysis for the behavioural sciences (3rd ed.). New York, NY: Routledge.
  3. ^ Hardy, Melissa (1993). Regression with dummy variables. Newbury Park, CA: Sage.

Further reading

[edit]
a货翡翠是什么意思 小孩便秘吃什么通便快 嗓子疼感冒吃什么药 精神小伙什么意思 灵芝孢子粉治什么病
为什么没人敢动景甜 五百年前是什么朝代 肚脐眼位置疼是什么原因 脸大剪什么发型好看 高血压是什么引起的
女孩叫锦什么好听 什么叫积阴德 什么屎不臭 孕囊小是什么原因 紫苏有什么作用
摸摸头是什么意思 缺铁性贫血吃什么 贪心不足蛇吞象什么意思 火是什么 腋下疣是什么原因造成的
胆汁反流用什么药好hcv8jop3ns3r.cn 腹部彩超可以检查什么huizhijixie.com 国企混改是什么意思hcv8jop2ns0r.cn 安络血又叫什么名hcv7jop7ns2r.cn 蜂蜡有什么用hcv9jop5ns6r.cn
男性阴囊瘙痒用什么药膏hcv9jop1ns7r.cn 刚愎自用什么意思hcv7jop5ns4r.cn 啤酒和什么不能一起吃hcv9jop0ns3r.cn 封建迷信是什么bfb118.com 彩字五行属什么hcv8jop9ns4r.cn
living是什么意思0297y7.com 牛蒡是什么hcv8jop5ns6r.cn 什么是高情商hcv8jop8ns6r.cn 眼睛像什么hcv9jop2ns6r.cn 白油是什么hcv8jop9ns3r.cn
孕妇贫血对胎儿有什么影响hcv8jop3ns4r.cn 脸霜什么牌子的好hcv8jop2ns5r.cn 什么食物对肺有好处hcv9jop1ns4r.cn 粉丝是什么做的inbungee.com sjh是什么意思hcv8jop4ns1r.cn
百度