Feature (machine learning)
This article needs more citations. (December 2014) |
| Part of a series on |
| Machine learning and data mining |
|---|
In machine learning and pattern recognition, a feature is an individual measurable property or characteristic of a data set.[1] Choosing informative, discriminating, and independent features is crucial to producing effective algorithms for pattern recognition, classification, and regression tasks. Features are usually numeric, but other types such as strings and graphs are used in syntactic pattern recognition, after some pre-processing step such as one-hot encoding.[2][3] The concept of "features" is related to that of explanatory variables used in statistical techniques such as linear regression.[1]
Feature types
[edit]In feature engineering, two types of features are commonly used: numerical and categorical.
Numerical features are continuous values that can be measured on a scale. Examples of numerical features include age, height, weight, and income. Numerical features can be used in machine learning algorithms directly.[4]
Categorical features are discrete values that can be grouped into categories. Examples of categorical features include gender, color, and zip code. Categorical features typically need to be converted to numerical features before they can be used in machine learning algorithms. This can be done using a variety of techniques, such as one-hot encoding, label encoding, and ordinal encoding.[5]
The type of feature that is used in feature engineering depends on the specific machine learning algorithm that is being used. Some machine learning algorithms, such as decision trees, can handle both numerical and categorical features. Other machine learning algorithms, such as linear regression, can only handle numerical features.
Classification
[edit]A numeric feature can be conveniently described by a feature vector. One way to achieve binary classification is using a linear predictor function (related to the perceptron) with a feature vector as input. The method consists of calculating the scalar product between the feature vector and a vector of weights, qualifying those observations whose result exceeds a threshold.[6]
Algorithms for classification from a feature vector include nearest neighbor classification, neural networks,[7] and statistical techniques such as Bayesian approaches.[8]
Examples
[edit]In character recognition, features may include histograms counting the number of black pixels along horizontal and vertical directions, number of internal holes, stroke detection and many others.[9]
In speech recognition, features for recognizing phonemes can include noise ratios, length of sounds, relative power, filter matches, logarithmic Mel-scale spectral vectors and Mel-frequency cepstral coefficients, which represent the frequency characteristics of audio signals.[10]
In spam detection algorithms, features may include the presence or absence of certain email headers, [11] the email structure, the language, the frequency of specific terms, the grammatical correctness of the text.[12][13]
In computer vision, there are a large number of possible features, such as edges and objects.[14]
Feature vectors
[edit]In pattern recognition and machine learning, a feature vector is an n-dimensional vector of numerical features that represent some object. Many algorithms in machine learning require a numerical representation of objects, since such representations facilitate processing and statistical analysis. When representing images, the feature values might correspond to the pixels of an image, while when representing texts the features might be the frequencies of occurrence of textual terms.[15] Feature vectors are equivalent to the vectors of explanatory variables used in statistical procedures such as linear regression.[16] Feature vectors are often combined with weights using a dot product in order to construct a linear predictor function that is used to determine a score for making a prediction.[6]
The vector space associated with these vectors is often called the feature space.[17] In order to reduce the dimensionality of the feature space, a number of dimensionality reduction techniques can be employed.[18]
Higher-level features can be obtained from already available features and added to the feature vector; for example, for the study of diseases the feature 'Age' is useful and is defined as Age = 'Year of death' minus 'Year of birth' . This process is referred to as feature construction.[19][20] Feature construction is the application of a set of constructive operators to a set of existing features resulting in construction of new features. Examples of such constructive operators include checking for the equality conditions {=, ≠}, the arithmetic operators {+,−,×, /}, the array operators {max(S), min(S), average(S)} as well as other more sophisticated operators, for example count(S, C)[21] that counts the number of features in the feature vector S satisfying some condition C or, for example, distances to other recognition classes generalized by some accepting device. Feature construction has long been considered a powerful tool for increasing both accuracy and understanding of structure, particularly in high-dimensional problems.[22] Applications include studies of disease and emotion recognition from speech.[23]
Selection and extraction
[edit]This section's style of writing may not reflect the encyclopedic tone used on Wikipedia. (August 2025) |
Raw features sets can be redundant or sufficiently large to make estimation and optimization difficult or ineffective. Therefore, many applications of machine learning and pattern recognition begin by selecting a subset of features or constructing a new and reduced set of features. These processes can facilitate learning and improve the generalization and interpretability of machine learning models.[24]
Feature selection and extraction involve evaluating different features and determining which features are most approprite for a particular machine learning task.Feature engineering refers to the process of developing and applying methods to select, transform, or construct features.[25] This process can involve automated techniques as well as the knowledge of the domain expert. Automating feature engineering can lead to feature learning, in which a machine learning system learns useful features from data rather than relying entirely on manually defined features.[26]
See also
[edit]References
[edit]- 1 2 Bishop, Christopher (2006). Pattern recognition and machine learning. Berlin: Springer. ISBN 0-387-31073-8.
- ↑ Flasiński, Mariusz; Jurek, Janusz (August 2014). "Fundamental methodological issues of syntactic pattern recognition". Pattern Analysis and Applications. 17 (3): 465–480. doi:10.1007/s10044-013-0322-1. ISSN 1433-7541.
- ↑ Yu, Lean; Zhou, Rongtian; Chen, Rongda; Lai, Kin Keung (2022-01-26). "Missing Data Preprocessing in Credit Classification: One-Hot Encoding or Imputation?". Emerging Markets Finance and Trade. 58 (2): 472–482. doi:10.1080/1540496X.2020.1825935. ISSN 1540-496X.
- ↑ "14. Feature Extraction — MGMT 4190/6560 Introduction to Machine Learning Applications @Rensselaer". introml.analyticsdojo.com. Retrieved 2026-09-09.
- ↑ Qiu, Qianxi; Liu, Han (2023-07-09). "Numerical Embedding of Categorical Features in Tabular Data: A Survey". 2023 International Conference on Machine Learning and Cybernetics (ICMLC). IEEE. pp. 446–451. doi:10.1109/ICMLC58545.2023.10327921. ISBN 979-8-3503-0378-0.
- 1 2 "Lecture 3: The Perceptron". www.cs.cornell.edu. Retrieved 2026-09-09.
- ↑ Mucherino, Antonio; Papajorgji, Petraq J.; Pardalos, Panos M. (2009), "k-Nearest Neighbor Classification", Data Mining in Agriculture, vol. 34, New York, NY: Springer New York, pp. 83–106, doi:10.1007/978-0-387-88615-2_4, ISBN 978-0-387-88614-5, retrieved 2026-09-09
{{citation}}: CS1 maint: work parameter with ISBN (link) - ↑ Lampinen, Jouko; Vehtari, Aki (April 2001). "Bayesian approach for neural networks—review and case studies". Neural Networks. 14 (3): 257–274. doi:10.1016/S0893-6080(00)00098-8. PMID 11341565.
- ↑ Govindan, V.K; Shivaprasad, A.P (January 1990). "Character recognition — A review". Pattern Recognition. 23 (7): 671–683. Bibcode:1990PatRe..23..671G. doi:10.1016/0031-3203(90)90091-X.
- ↑ Jurafsky, Daniel; Martin, James H. "Speech and Language Processing (3rd ed. draft), Chapter 14: Speech Recognition" (PDF). Stanford University. Retrieved 2026-04-15.
- ↑ Salcedo-Campos, Francisco; Díaz-Verdejo, Jesús; García-Teodoro, Pedro (July 2012). "Segmental parameterisation and statistical modelling of e-mail headers for spam detection". Information Sciences. 195: 45–61. doi:10.1016/j.ins.2012.01.022.
- ↑ Williams, Sarah E.; Sarno, Dawn M.; Lewis, Joanna E.; Shoss, Mindy K.; Neider, Mark B.; Bohil, Corey J. (2019-08-03). "The psychological interaction of spam email features". Ergonomics. 62 (8): 983–994. doi:10.1080/00140139.2019.1614681. ISSN 0014-0139. PMC 6629481. PMID 31056018.
- ↑ Zhang, Yudong; Wang, Shuihua; Phillips, Preetha; Ji, Genlin (July 2014). "Binary PSO with mutation operator for feature selection using decision tree applied to spam detection". Knowledge-Based Systems. 64: 22–31. doi:10.1016/j.knosys.2014.03.015.
- ↑ "Features". CIRL. 2013-09-16. Retrieved 2026-09-09.
- ↑ "Lecture 1: Supervised Learning". www.cs.cornell.edu. Retrieved 2026-09-09.
- ↑ Wittek, Peter (2014), "Quantum Computing", Quantum Machine Learning, Elsevier, pp. 11–24, doi:10.1016/b978-0-12-800953-6.00004-9, ISBN 978-0-12-800953-6, retrieved 2026-09-09
{{citation}}: CS1 maint: work parameter with ISBN (link) - ↑ Bergmann, Dave (2025-01-28). "What Is Latent Space? | IBM". www.ibm.com. Retrieved 2026-09-09.
- ↑ Jia, Weikuan; Sun, Meili; Lian, Jian; Hou, Sujuan (June 2022). "Feature dimensionality reduction: a review". Complex & Intelligent Systems. 8 (3): 2663–2693. doi:10.1007/s40747-021-00637-x. ISSN 2199-4536.
- ↑ Liu, H., Motoda H. (1998) Feature Selection for Knowledge Discovery and Data Mining., Kluwer Academic Publishers. Norwell, MA, USA. 1998.
- ↑ Piramuthu, S., Sikora R. T. Iterative feature construction for improving inductive learning algorithms. In Journal of Expert Systems with Applications. Vol. 36, Iss. 2 (March 2009), pp. 3401-3406, 2009
- ↑ Bloedorn, E., Michalski, R. Data-driven constructive induction: a methodology and its applications. IEEE Intelligent Systems, Special issue on Feature Transformation and Subset Selection, pp. 30-37, March/April, 1998
- ↑ Breiman, L. Friedman, T., Olshen, R., Stone, C. (1984) Classification and regression trees, Wadsworth
- ↑ Sidorova, J., Badia T. Syntactic learning for ESEDA.1, tool for enhanced speech emotion detection and analysis. Internet Technology and Secured Transactions Conference 2009 (ICITST-2009), London, November 9–12. IEEE
- ↑ Hastie, Trevor; Tibshirani, Robert; Friedman, Jerome H. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer. ISBN 978-0-387-84884-6.
- ↑ Duboue, Pablo, ed. (2020), "Features, Reduced: Feature Selection, Dimensionality Reduction and Embeddings", The Art of Feature Engineering: Essentials for Machine Learning, Cambridge: Cambridge University Press, pp. 79–111, doi:10.1017/9781108671682.006, ISBN 978-1-108-70938-5, retrieved 2026-09-08
{{citation}}: CS1 maint: work parameter with ISBN (link) - ↑ Duboue, Pablo, ed. (2020), "Advanced Topics: Variable-Length Data and Automated Feature Engineering", The Art of Feature Engineering: Essentials for Machine Learning, Cambridge: Cambridge University Press, pp. 112–136, doi:10.1017/9781108671682.007, ISBN 978-1-108-70938-5, retrieved 2026-09-08
{{citation}}: CS1 maint: work parameter with ISBN (link)