Lets explore the stack
published at 02.10.2026 16:56 by Jens Weller
Save to Instapaper Pocket
Part two in the new C++ content, as I've spend some time creating content to be now released in front of Meeting C++ 2026.
Maybe read last weeks post About alignment, struct layout and the cache first.
Maybe its a good idea to continue in an area which a lot of other things build upon: the stack, followed by the heap and visiting RAII. Modern programs usually have other memory areas too, like read only sections or a place where static text is stored. Exploring these is also fun, but gets one into an area ruled by the OS and hardware dependent memory layout. Which may differ between the wide range of C++ usage. This blog entry does a good job in going into some of these details.
I got curious about the stack once I wondered why there is such a difference between stack and heap, yet they don’t really exist in hardware. There is no special cache or memory for stack or heap in our processors or GPUs. I did also wonder which tools/features existed for either.
But first, lets quickly recall that C++ runs on an abstract machine. There is no official implementation, the committee produces a (theoretical) standard only. Today there are 3 main implementations (clang, GCC, MSVC) and a few others. And often the standard leaves the details to the implementation, defining things as implementation defined. In the words of the standard C++ is running on an abstract machine. Implementations of C++ then have to fill the gap between being abstract and the targeted platform.
The stack exists in memory during program execution, when a function is called, a stack frame is allocated and initialized with its various variables, potential parameters. This frame is deallocated once the function returns, when a function calls it self a new stack frame is allocated for this call. That is why recursion has its limits, and can exhaust the available memory space assigned for the call stack. With constexpr ofc we now also have stacks at compile time, but that is a different story.
For exploring the actual stack space of a function, C++ it self does not offer such tools. GCC and clang have a build in __builtin_frame_address, which returns a void pointer to the current stack frame start. With this one can toy around a little bit to see how the distance to the last variable in a function changes when adding various variables or control structures to the body of the function.
The __builtin_frame_address is part of a GCC API to allow for exploring the callers of a function, __buildin_frame_address(0) will get you the frame address of this current function, while higher values (1,2…) go up the call stack or “Calling this function with a nonzero argument can have unpredictable effects, including crashing the calling program”.
A simple function exploring this:
void check_stack()
{
// Get address of current stack frame (depth 0 = current function)
void* frame_address = __builtin_frame_address(0);
size_t stack_used=0;
int x=0;
void* pend = static_cast<void*>(&x);
stack_used = ((char*)frame_ address - (char*)pend) + sizeof(pend);
std::cout << “Stack used: "<<stack_used<<" bytes\n";
}
There also exists -fstack-usage which writes a .su file for each .cpp file containing a listing of function names and their stack sizes. If you want to explore this a bit further, you can do so in Compiler Explorer. The stack viewer feature seems to only work on clang currently though.
A fun exercise is to create a small function returning a random int, you can put things like random_device or a mersenne_twister on the stack, and see how much memory they take up. Making a variable static removes this footprint from the stack towards a block of memory of objects in static storage initialized with the first function call. This is since C++11 also thread safe. A known very similar pattern is the Meyers singleton, where you declare the singleton as a static variable in the function returning the singleton.
But writing the code showing struct layout raised the question for me if also variables on the stack of a function are aligned and if this can be shown. Its not as easy as with a struct, as a variable could be placed into a register – but taking the address of a variable to know its position in memory should also prevent this from occurring.
Adopting above’s code to also show the padding bytes sounded trivial. Reuse the code for the struct padding and take each address of the variables I’d like to show in the memory map. This seemed to work at first, until I’ve noticed that it placed the next variable inside the previous bigger one. The addressing on the stack does work different then one would assume when learning this from how struct layout works.
For these variables on the stack:
int s =0; // 0 size_t stack_used = 0; // 1 short c = 42; // 2 int x=0+c+s; // 3 void* pend = static_cast<void*>(&x); // 4
Lets print out the size and memory addresses, printing the address as an actual number, not a hexadecimal 0x value:
auto print_info = [](std::string_view name, auto& t){
std::cout << name << ": address: "<<(unsigned long)&t<<" size: "<<sizeof(t) << "\n"; };
print_info("s ",s);
print_info("su",stack_used);
print_info("c ",c);
print_info("x ",x);
print_info("pe",pend);
Which gives these results:
s : address: 140730290274324 size: 4 // int
su: address: 140730290274312 size: 8 // size_t stack_size
c : address: 140730290274310 size: 2 // short
x : address: 140730290274304 size: 4 // int
pe: address: 140730290274296 size: 8 // void* pend
This shows that s is the largest address with ...24, with the other addresses being placed at a lower address in memory. To the second variable is a distance of 12 bytes, but only 8 bytes are the size of the type, so padding bytes must exist here. The short after stack_size (the second variable) is placed two bytes after it, which also indicates that the 8 bytes of stack_size should be between 12 and 20. This also means that the first variable s lives at bytes 24 – 28. That brings up the question, does one want to show the stack variables in order of declaration or how they are layed out in memory?
In the order of declaration one gets this memory map, with -1 being a padding byte:
GCC 15.2: [0, 0, 0, 0, -1, -1, -1, -1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, -1, -1, 3, 3, 3, 3, 4, 4, 4, 4, 4, 4, 4, 4]
clang 21.1: [0, 0, 0, 0, -1, -1, -1, -1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, -1, -1, 3, 3, 3, 3, 4, 4, 4, 4, 4, 4, 4, 4]
ARM64 GCC 15.2 [0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, -1, -1, 3, 3, 3, 3, 4, 4, 4, 4, 4, 4, 4, 4]
Which means that the basic rules for achieving smaller sizes for structs to avoid excess padding can also be applied to your functions! It may make the size of your function as a stack frame smaller. But the impact is far smaller then for data structures. A function is not a struct, and its size is not only defined by its stack variables in its code block. While you can show that a function with the same variables gets bigger if you declare them in an order that creates a lot of padding bytes, it was not enough to influence the performance in my tests. And the way the code works prevents any placement of these variables in registers, which is a common thing other wise. So its nice to see that padding bytes also exist in functions, but one needs to take into account that this does not show the memory layout of common functions using registers.
One thing I did learned through this is, that accessing the various variables *seems* to be faster when you do this in order of declaration. This caused a minor difference between the two functions, and then the longer function is faster – as its accessing the variables in declaration order. You can play around with this in quick-bench if you like.
Sharing the code that generates above’s output with you, the refactored function from the previous code example:
void check_stack() {
// Get address of current stack frame (depth 0 = current function)
void* frame_address = __builtin_frame_address(0);
//some test variables to create a certain layout to show padding bytes
int s =0;
size_t stack_used = 0;
short c ='.';
int x=0+c+s;
void* pend = static_cast<void*>(&x);
stack_used = ((char*)frame_address - (char*)pend) + sizeof(pend);
std::vector stack(((char*)&s - (char*)&pend),-1);
size_t offset = sizeof(s);
std::vector stack(offset+((char*)&s - (char*)&pend),-1);
write_to_array(&stack[0],offset + ((char*)&s - (char*)&s),sizeof(int),0 );
write_to_array(&stack[0],offset +((char*)&s - (char*)&stack_used),sizeof(size_t),1 );
write_to_array(&stack[0],offset +((char*)&s - (char*)&c),sizeof(short),2 );
write_to_array(&stack[0],offset +((char*)&s - (char*)&x),sizeof(int),3 );
write_to_array(&stack[0],offset +((char*)&s - (char*)&pend),sizeof(void*),4 );
std::println("{}",stack);
//showing memory positions of the variables
std::cout << &stack_used << "\n" << &c << "\n";
}
Added to the function check_stack has been a block that writes the positions into the vector calling a helper function write_to_array. The argument for the end of the variables memory block needs to be calculated from the distance of its position to the first variable, plus the size in bytes of that variable. The result is then a memory map showing the reversed layout of the actual memory, with the first variable at the beginning. To show the layout how it is in memory, one needs to reverse the memory map vector, showing the actual memory layout:
[4, 4, 4, 4, 4, 4, 4, 4, 3, 3, 3, 3, -1, -1, 2, 2, 1, 1, 1, 1, 1, 1, 1, 1, -1, -1, -1, -1, 0, 0, 0, 0]
4 is void* pend, 3 is int x, -1 padding, 2 is short c, 1 size_t stack_used, padding followed by 1 - int s.
The code for write_to_array:
template< class Type>
constexpr void write_to_array(Type* arr,size_t end, size_t size, Type value)
{
size_t start = end - size;
std::println("start:{},end:{}",start,end);
do
{
// if(arr[start]==-1)
arr[start]=val;
}while(++start < end);
}
The function write_to_array takes now the pointer to the first element of the array, the end position in the array, the size of the type in bytes and the value to write to the array from start to end. It needs to calculate the start from end and size and then do a loop.
The code on Compiler Explorer also includes the two tested functions from the mentioned test on quick-bench, as this shows their stack-size.
The memory on the stack is a bit outside your control – but in a good way, and may involve details such as the calling convention and ABI your code compiles to. While the life time of an object is defined by its scope, its stack memory is freed when the memory of the stack frame is cleared, not when the scope is left. Placing objects on the stack brings the advantage that the memory gets loaded and partially initialized before the function runs. Your variables are already in cache, while a value on the heap accessed through a pointer may cause an indirection and an additional load. It is likely that it too is getting pre-fetched, if it exists when the function is called.
One could also go deeper and take a look at the behavior of stack variables in if/else blocks or in loops. I leave that to the reader, as I think its not adding much to your C++ knowledge. Also above’s code likely shows platform specific behavior, as the C++ standard runs on an abstract machine. Hence such details are often implementation defined.
Join the Meeting C++ patreon community!
This and other posts on Meeting C++ are enabled by my supporters on patreon!