C11은 C99 표준을 대체했고, C18은 C11을 대체했다. C18은 새로운 기능을 제공하지 않고 C11의 결함을 수정했을 뿐이다.
모든 C 표준은 어느 버전을 사용하는지 알아낼 수 있는 특별한 매크로를 정의한다. 상이한 버전의 C 표준을 지원하는 컴파일러 사용 시 가장 먼저 버전을 확인해야 한다. gcc는 4.7 버전부터 C11을 지원한다.
코드 박스 12-1 [예제 12-1] C 표준 버전을 감지하기
#include <stdio.h>
int main(int argc, char** argv) {
#if __STDC_VERSION__ >= 201710L
printf("Hello World from C18!\n");
#elif __STDC_VERSION__ >= 201112L
printf("Hello World from C11!\n");
#elif _STDC_VERSION__ >= 199001L
printf("Hello World from C99!\n");
#else
printf("Hello World from C89/C90!\n");
#endif
return 0;
}
컴파일러에 -std=CXX 옵션을 전달하여 특정 버전의 C 표준을 사용하도록 요청할 수 있다.
셀 박스 12-1 다양한 버전의 C 표준으로 [예제 12-1] 컴파일하기
$ gcc 12_1.c -o 12_1.out
$ ./12_1.out
Hello World from C18!
$ gcc 12_1.c -o 12_1.out -std=c11
$ ./12_1.out
Hello World from C11!
$ gcc 12_1.c -o 12_1.out -std=c99
$ ./12_1.out
Hello World from C89/C90!
$ gcc 12_1.c -o 12_1.out -std=c90
$ ./12_1.out
Hello World from C89/C90!
$ gcc 12_1.c -o 12_1.out -std=c89
$ ./12_1.out
Hello World from C89/C90!
최신 컴파일러의 기본 C 표준 버전은 C18이다.
C11에서는 버퍼 오버플로 공격의 대상이 되는 gets 함수가 제거되었다. 대신 fgets 함수를 사용하도록 강제된다.
일반적으로 fopen 함수는 파일을 열어 그 디스크립터를 반환할 때 사용한다.
코드 박스 12-2 fopen 함수 패밀리에 대한 다양한 시그니처
FILE *fopen(const char *restrict path, const char *restrict mode);
FILE *fdopen(int fd, const char *mode);
FILE *freopen(const char *restrict path, const char *restrict mode, FILE *restrict stream);
모든 시그니처는 파일을 여는 방식에 관한 문자열인 mode를 매개변수로 정의한다.
셀 박스 12-2 Ubuntu의 fopen 매뉴얼 페이지에서 발췌
DESCRIPTION
The fopen() function opens the file whose name is the string pointed to by path and associates a stream
with it.
The argument mode points to a string beginning with one of the following sequences (possibly followed by
additional characters, as described below):
r Open text file for reading. The stream is positioned at the beginning of the file.
r+ Open for reading and writing. The stream is positioned at the beginning of the file.
w Truncate file to zero length or create text file for writing. The stream is positioned at the be‐
ginning of the file.
w+ Open for reading and writing. The file is created if it does not exist, otherwise it is truncated.
The stream is positioned at the beginning of the file.
a Open for appending (writing at end of file). The file is created if it does not exist. The stream
is positioned at the end of the file.
a+ Open for reading and appending (writing at end of file). The file is created if it does not exist.
Output is always appended to the end of the file. POSIX is silent on what the initial read posi‐
tion is when using this mode. For glibc, the initial file position for reading is at the beginning
of the file, but for Android/BSD/MacOS, the initial file position for reading is at the end of the
file.
┌──────────────┬───────────────────────────────┐
│ fopen() mode │ open() flags │
├──────────────┼───────────────────────────────┤
│ r │ O_RDONLY │
├──────────────┼───────────────────────────────┤
│ w │ O_WRONLY | O_CREAT | O_TRUNC │
├──────────────┼───────────────────────────────┤
│ a │ O_WRONLY | O_CREAT | O_APPEND │
├──────────────┼───────────────────────────────┤
│ r+ │ O_RDWR │
├──────────────┼───────────────────────────────┤
│ w+ │ O_RDWR | O_CREAT | O_TRUNC │
├──────────────┼───────────────────────────────┤
│ a+ │ O_RDWR | O_CREAT | O_APPEND │
└──────────────┴───────────────────────────────┘
NOTES
glibc notes
The GNU C library allows the following extensions for the string specified in mode:
c (since glibc 2.3.3)
Do not make the open operation, or subsequent read and write operations, thread cancelation points.
This flag is ignored for fdopen().
e (since glibc 2.7)
Open the file with the O_CLOEXEC flag. See open(2) for more information. This flag is ignored for
fdopen().
m (since glibc 2.3)
Attempt to access the file using mmap(2), rather than I/O system calls (read(2), write(2)). Cur‐
rently, use of mmap(2) is attempted only for a file opened for reading.
x Open the file exclusively (like the O_EXCL flag of open(2)). If the file already exists, fopen()
fails, and sets errno to EEXIST. This flag is ignored for fdopen().
In addition to the above characters, fopen() and freopen() support the following syntax in mode:
,ccs=string
The given string is taken as the name of a coded character set and the stream is marked as wide-oriented.
Thereafter, internal conversion functions convert I/O to and from the character set string. If the
,ccs=string syntax is not specified, then the wide-orientation of the stream is determined by the first
file operation. If that operation is a wide-character operation, the stream is marked wide-oriented, and
functions to convert to the coded character set are loaded.
w 또는 w+ 모드는 이미 파일이 존재할 시 삭제(truncate)하는 것이 기본 동작인데, 이로 인한 파일 삭제를 방지하기 위해 기존에는 stat과 같은 파일 시스템 API를 사용해 파일이 존재하는지 여부를 검증하거나 a 모드를 사용했다. C11에서 도입된 x 모드 덕분에 이러한 상용구가 축소되었다.
또한 C11에서는 버퍼의 경계 검사를 수행하는 fopen_s API도 도입되었다.
버퍼 경계를 넘어선 접근은 버퍼 오버플로를 발생시킨다. 이러한 유형의 C 프로그램 악용(exploitation)은 서비스 거부(DoS) 공격에 활용되기도 한다. 대표적으로 strcpy, strcat과 같은 string.h 헤더에 정의된 문자열 조작 함수는 버퍼 오버플로 공격 방지를 위한 경계 검사 메커니즘이 결여되어 있다.
C11에서는 _s 접미사가 붙는 경계 검사 함수가 추가되었다. strcpy_s, strcat_s과 같은 함수는 경계 검사 메커니즘을 갖는다.
코드 박스 12-3 strcpy_s 함수의 시그니처
errno_t strcpy_s(char *restrict dest, rsize_t destsz, const char *restrict src);
코드 박스 12-4 값을 반환하지 않는 함수의 예
void main_loop() {
while (1) {
...
}
}
int main(int argc, char** argv) {
...
main_loop();
return 0;
}
main_loop 함수는 프로그램의 주요 작업을 수행하며, 함수에서 반환되면 프로그램도 종료되는 것으로 간주된다. 이러한 예외적인 경우에 컴파일러는 더 많은 최적화를 수행할 수 있다. 하지만 main_loop 함수가 절대 반환되지 않는다는 점을 알아야 한다.
C11에서는 이처럼 값을 반환하지 않는 함수는 stdnoreturn.h 헤더에 정의된 _Noreturn 키워드로 지정할 수 있다.
코드 박스 12-5 main_loop를 끝나지 않는 함수로 표시하는 _Noreturn 키워드 사용하기
_Noreturn void main_loop() {
while (true) {
...
}
}
프로그램의 빠른 종료를 위해 C11에서 추가된 exit, quick_exit, abort 같은 함수들은 값을 반환하지 않는 함수로 간주되어 컴파일러가 인식하고 적절한 경고를 표시할 수 있다. Noreturn으로 표시한 함수가 반환 값이 있는 경우는 정의되지 않는 행위이다.
C11에서 추가된 _Generic 키워드는 컴파일 중 자료형을 인식하는 매크로를 작성하는 데 사용할 수 있다. 보편적으로 일반 선택(generic selection)이라 한다.
코드 박스 12-6 [예제 12-2] 제네릭 매크로의 예
#include <stdio.h>
#define abs(x) _Generic((x), \
int: absi, \
double: absd)(x)
int absi(int a) {
return a > 0 ? a : -a;
}
double absd(double a) {
return a > 0 ? a : -a;
}
int main(int argc, char** argv) {
printf("abs(-2): %d\n", abs(-2));
printf("abs(2.5): %f\n", abs(2.5));
return 0;
}
x 인수의 자료형에 따라 int 값이면 absi, double 값이면 absd를 사용하도록 한다. 이는 이전의 C 컴파일러에서도 확인할 수 있지만 C11에서 표준이 되었다.
C11 표준에는 UTF-8, UTF-16, UTF-32 인코딩을 통한 유니코드 지원이 추가되었다. 이전의 C에서는 IBM 유니코드 국제 컴포넌트(ICU: IBM International Components for Unicode)와 같은 서드 파티 라이브러리를 사용해야 했다.
C11 이전에는 아스키(ASCII)와 확장 아스키(Extended ASCII) 문자를 저장하는 8비트 변수인 char, unsigned char 형만 있었다. C11부터는 새로운 문자에 대한 지원이 추가되어 각 문자가 1바이트를 초과하는 확장 문자(wide character) 기반의 문자열 사용이 가능해졌다.
유니코드 표준은 1바이트 이상을 사용해 아스키, 확장 아스키, 확장 문자에 있는 모든 문자를 인코딩하는 여러 인코딩(encoding) 방법을 도입했다.
UTF-8은 일반적으로 한 문자가 1~4바이트의 크기를 가진다. 비트 패턴은 각각 1바이트면 0xxxxxxx, 2바이트면 110xxxxx 10xxxxxx, 3바이트면 1110xxxx 10xxxxxx 10xxxxxx, 4바이트면 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx와 같다. 따라서 가변 길이 인코딩(variable-sized encoding)으로 간주되며, 아스키의 상위 집합(superset)으로 간주된다.
UTF-16은 2바이트를 기본 단위로 삼아, U+0000~U+FFFF 범위의 코드포인트는 2바이트로 표현하고, 그보다 큰 범위의 코드포인트는 4바이트로 표현한다. 따라서 이 역시 가변 길이 인코딩으로 간주된다.
UTF-32는 모든 문자에 대한 값을 저장하기 위해 오직 4바이트만 사용한다. 그러므로 고정 길이 인코딩으로 간주된다. 아스키 문자에도 고정된 4바이트를 사용하기 때문에 더 많은 메모리 공간을 소비하지만, 압축 인코딩 방식인 UTF-8, UTF-16과 달리 문자의 실제 값을 반환하기 위한 연산 횟수가 절약된다는 장점이 있다.
코드 박스 12-7 [예제 12-3]을 빌드하는 데 필요한 헤더 파일과 선언
#include <stdlib.h>
#include <stdio.h>
#include <string.h>
#ifdef __APPLE__
#include <stdint.h>
typedef uint16_t char16_t;
typedef uint32_t char32_t;
#else
#include <uchar.h> // needed for char16_t and char32_t
#endif
C11에서는 유니코드 문자열에서 작동하는 유틸리티 함수를 제공하지 않는다. 따라서 strlen을 새로 작성해야 한다.
코드 박스 12-8 [예제 12-3]에서 사용된 함수의 정의
typedef struct {
long num_chars;
long num_bytes;
} unicode_len_t;
unicode_len_t strlen_ascii(char* str) {
unicode_len_t res;
res.num_chars = 0;
res.num_bytes = 0;
if (!str) {
return res;
}
res.num_chars = strlen(str) + 1;
res.num_bytes = strlen(str) + 1;
return res;
}
unicode_len_t strlen_u8(char* str) {
unicode_len_t res;
res.num_chars = 0;
res.num_bytes = 0;
if (!str) {
return res;
}
// the last null character
res.num_chars = 1;
res.num_bytes = 1;
while (*str) {
unsigned char c = (unsigned char) *str;
if ((c & 0x80) == 0x00) { // 0xxxxxxx - 1 byte
++res.num_chars;
++res.num_bytes;
++str;
} else if ((c & 0xe0) == 0xc0) { // 110xxxxx - 2 bytes
++res.num_chars;
res.num_bytes += 2;
str += 2;
} else if ((c & 0xf0) == 0xe0) { // 1110xxxx - 3 bytes
++res.num_chars;
res.num_bytes += 3;
str += 3;
} else if ((c & 0xf8) == 0xf0) { // 11110xxx - 4 bytes
++res.num_chars;
res.num_bytes += 4;
str += 4;
} else {
fprintf(stderr, "UTF-8 string is not valid!\n");
exit(1);
}
}
return res;
}
unicode_len_t strlen_u16(char16_t* str) {
unicode_len_t res;
res.num_chars = 0;
res.num_bytes = 0;
if (!str) {
return res;
}
// the last null character
res.num_chars = 1;
res.num_bytes = 2;
while (*str) {
if (*str < 0xdc00 || *str > 0xdfff) {
++res.num_chars;
res.num_bytes += 2;
++str;
} else {
++res.num_chars;
res.num_bytes += 4;
str += 2;
}
}
return res;
}
unicode_len_t strlen_u32(char32_t* str) {
unicode_len_t res;
res.num_chars = 0;
res.num_bytes = 0;
if (!str) {
return res;
}
// the last null character
res.num_chars = 1;
res.num_bytes = 4;
while (*str) {
++res.num_chars;
res.num_bytes += 4;
++str;
}
return res;
}
코드 박스 12-9 [예제 12-3]의 main 함수
int main(int argc, char** argv) {
char ascii_string[32] = "Hello World!";
char utf8_string[32] = u8"Hello World!";
char utf8_string_2[32] = u8"درود دنیا!";
char16_t utf16_string[32] = u"Hello World!";
char16_t utf16_string_2[32] = u"درود دنیا!";
char16_t utf16_string_3[32] = u"𝦵হ𝡩!";
char32_t utf32_string[32] = U"Hello World!";
char32_t utf32_string_2[32] = U"درود دنیا!";
char32_t utf32_string_3[32] = U"𝦵হ𝡩!";
unicode_len_t len = strlen_ascii(ascii_string);
printf("Length of ASCII string:\t\t\t %ld chars, %ld bytes\n\n",
len.num_chars, len.num_bytes);
len = strlen_u8(utf8_string);
printf("Length of UTF-8 english string:\t\t %ld chars, %ld bytes\n",
len.num_chars, len.num_bytes);
len = strlen_u16(utf16_string);
printf("Length of UTF-16 english string:\t %ld chars, %ld bytes\n",
len.num_chars, len.num_bytes);
len = strlen_u32(utf32_string);
printf("Length of UTF-32 english string:\t %ld chars, %ld bytes\n\n",
len.num_chars, len.num_bytes);
len = strlen_u8(utf8_string_2);
printf("Length of UTF-8 persian string:\t\t %ld chars, %ld bytes\n",
len.num_chars, len.num_bytes);
len = strlen_u16(utf16_string_2);
printf("Length of UTF-16 persian string:\t %ld chars, %ld bytes\n",
len.num_chars, len.num_bytes);
len = strlen_u32(utf32_string_2);
printf("Length of UTF-32 persian string:\t %ld chars, %ld bytes\n\n",
len.num_chars, len.num_bytes);
len = strlen_u16(utf16_string_3);
printf("Length of UTF-16 alien string:\t\t %ld chars, %ld bytes\n",
len.num_chars, len.num_bytes);
len = strlen_u32(utf32_string_3);
printf("Length of UTF-32 alien string:\t\t %ld chars, %ld bytes\n",
len.num_chars, len.num_bytes);
return 0;
}
이 예제는 C11 이상의 컴파일러로만 컴파일할 수 있다.
셀 박스 12-3 [예제 12-3]을 컴파일하고 실행하기
$ gcc 12_3.c -std=c18 -o 12_3.out
$ ./12_3.out
Length of ASCII string: 13 chars, 13 bytes
Length of UTF-8 english string: 13 chars, 13 bytes
Length of UTF-16 english string: 13 chars, 26 bytes
Length of UTF-32 english string: 13 chars, 52 bytes
Length of UTF-8 persian string: 11 chars, 19 bytes
Length of UTF-16 persian string: 11 chars, 22 bytes
Length of UTF-32 persian string: 11 chars, 44 bytes
Length of UTF-16 alien string: 5 chars, 14 bytes
Length of UTF-32 alien string: 5 chars, 20 bytes
아스키 문자는 UTF-8 체계에서 1바이트만 사용하기 때문에 대부분의 문자가 아스키 문자일 경우 효율적이다. 아시아 언어에서 사용하는 문자를 대부분 사용한다면, 대부분의 문자가 최대 2바이트까지 사용하는 UTF-16이 효율적이다. UTF-32는 거의 사용되지 않지만 고정 길이의 코드 인쇄(code print)가 유용한 시스템에서 사용할 수 있다. 특히 데이터 인덱스를 구축하거나 일부 병렬 처리 파이프라인(parallel processing pipeline)에서 이득이 있을 수 있다.
익명 구조체와 익명 공용체는 보통 다른 자료형 내에서 중첩 자료형으로 사용된다.
코드 박스 12-10 [예제 12-4] 익명 구조체와 익명 공용체가 함께 있는 예제
typedef struct {
union {
struct {
int x;
int y;
};
int data[2];
};
} point_t;
이 자료형은 익명 구조체와 바이트 배열 필드 데이터가 같은 메모리 공간을 공유한다.
코드 박스 12-11 [예제 12-4] 익명 구조체와 함께 익명 공용체를 사용하는 main 함수
#include <stdio.h>
typedef struct {
union {
struct {
int x;
int y;
};
int data[2];
};
} point_t;
int main(int argc, char** argv) {
point_t p;
p.x = 10;
p.data[1] = -5;
printf("Point (%d, %d) using an anonymous structure inside an anonymous union.\n", p.x, p.y);
printf("Point (%d, %d) using byte array inside an anonymous union.\n", p.data[0], p.data[1]);
return 0;
}
셀 박스 12-4 [예제 12-4]를 컴파일하고 실행하기
$ gcc 12_4.c -std=c11 -o 12_4.out
$ ./12_4.out
Point (10, -5) using an anonymous structure inside an anonymous union.
Point (10, -5) using byte array inside an anonymous union.
C에서 멀티스레딩은 POSIX API 또는 pthread 라이브러리를 통해 오랫동안 지원되었다. POSIX API는 리눅스나 유닉스 계열의 POSIX 호환 시스템에서만 사용할 수 있다. 이외의 시스템은 운영체제가 제공하는 라이브러리를 사용해야 한다. C11에서는 POSIX 호환과 상관없이 표준 C를 사용하는 모든 시스템에서 사용할 수 있는 표준 스레딩 라이브러리를 제공한다.