R Advanced

is.na() Function: Checking for Missing Data in R

The is.na() function checks for missing values (NA) in the R object. It returns TRUE for NA values and FALSE otherwise. The valid object can be anything like a data frame, matrix, list, or vector.

It is extremely helpful in data cleaning and preparation, as it helps identify and handle missing values in a dataset.

Syntax

is.na(obj)

Parameters

Name Value
obj It is an input object that needs to be tested for NA value. The object can be anything from a vector, list, matrix, or data frame.

With DataFrame

If you have a data frame and you are not sure how many NA values are there in the data frame, you can use the is.na() function and pass the data frame will return a data frame where NA values are replaced by TRUE and, in another case, FALSE.

df <- data.frame(
  col1 = c(1, NA, 3),
  col2 = c(NA, 5, NA),
  col3 = c(7, NA, 9)
)

is.na(df)

Output

With Vector

From the visual representation, you can see that we created a vector with two NA values and use is.na() function that will return TRUE for NA values and FALSE otherwise.

vec <- c(11, 21, 19, NA, 46, NA)

is.na(vec)

Output

[1] FALSE  FALSE  FALSE  TRUE  FALSE  TRUE

Finding positions of NAs

The any() function returns whether any values are NA in the input object.

data <- c(11, 21, 19, NA, 46, NA)

any(is.na(data))

Output

[1] TRUE

In this example, any() function returns TRUE because the vector data contains at least one NA value. If it does not have a single NA value, then it returns FALSE.

data <- c(11, 21, 19, 46, 18)

any(is.na(data))

Output

[1] FALSE

Counting NA values in a data frame

When doing exploratory data analysis, finding and removing NA values is the most important part; these functions will help you find them.

You can count total NA values in a data frame by combining is.na() and sum() functions.

Let’s take an example data frame df and count the NA values.

df <- data.frame(
  col1 = c(1, NA, 3),
  col2 = c(NA, 5, NA),
  col3 = c(7, NA, 9)
)

num_na_df <- sum(is.na(df))

num_na_df

Output

[1] 4

Counting NA values in a vector

You can count the number of NA values in a vector using the combination of sum() and is.na() functions.

vec <- c(11, 21, 19, NA, 46, NA)

sum(is.na(vec))

Output

[1] 11  21  19  46

To deal with NA values, you might use functions like na.omit() to remove rows with NA or functions like replace(), mean(), median(), etc., to impute missing values.

Removing NAs (be cautious!)

We can remove the NA values from a vector using the “!” operator and is.na() function.

vec <- c(11, 21, 19, NA, 46, NA)

vec[!is.na(vec)]

Output

[1] 11  21  19  46

You can see from the above output that we removed NA values from the vector.

Recent Posts

cbind() Function: Binding R Objects by Columns

R cbind (column bind) is a function that combines specified vectors, matrices, or data frames…

1 week ago

rbind() Function: Binding Rows in R

The rbind() function combines R objects, such as vectors, matrices, or data frames, by rows.…

1 week ago

as.numeric(): Converting to Numeric Values in R

The as.numeric() function in R converts valid non-numeric data into numeric data. What do I…

2 weeks ago

Calculating Natural Log using log() Function in R

The log() function calculates the natural logarithm (base e) of a numeric vector. By default,…

3 weeks ago

Dollar Sign ($ Operator) in R

In R, you can use the dollar sign ($ operator)  to access elements (columns) of…

1 month ago

Calculating Absolute Value using abs() Function in R

The abs() function calculates the absolute value of a numeric input, returning a non-negative (only…

1 month ago