rapidyaml
0.16.0
parse and emit YAML, and do it fast
Loading...
Searching...
No Matches
doxy_serialization_user_types.hpp
Go to the documentation of this file.
1
2
// DANGER: Keep markdown []() links in a single line!!!
3
//
4
// doxygen is broken and fails to render the markdown links when
5
// they span multi lines.
6
7
8
#include "
c4/yml/tree.hpp
"
9
#include "
c4/yml/node.hpp
"
10
#include "
c4/yml/scalar_charconv.hpp
"
11
12
namespace
c4
{
13
namespace
yml
{
14
15
16
/** @addtogroup doc_serialization_user_types
17
18
<br>
19
<hr>
20
## Serialization type categories
21
22
There are two distinct type categories to consider regarding YAML
23
serialization:
24
25
- **Container types**. These represent a hierarchy of values (or
26
containers) and must converted to/from a YAML map (@ref MAP) or
27
sequence (@ref SEQ).
28
29
- **Scalar types**. These types are encoded as scalars, but need
30
to be transformed from/to their string representation in the
31
YAML buffer.
32
33
34
A container type will always require child nodes in the tree. A scalar
35
type will always be a leaf (childless) node in the tree. Most of the
36
time, a scalar will be converted to string and not require any meta
37
info (like tags) or style flags set in the tree, but occasionally this
38
will be needed.
39
40
So in fact, from the implementation point of view, the categories are
41
the following:
42
43
- **General types**. Require extra structure/info from the tree:
44
child nodes (required by containers) and/or tags or extra @ref
45
NodeType flags (required by some scalars).
46
47
- **Scalar types**. These merely need to be converted to string and
48
then set as scalars on the tree, without needing to set any tags
49
or extra @ref NodeType flags.
50
51
52
To have rapidyaml interact with your types, you need to define functions
53
where this is done, and then the compiler will have rapidyaml call your
54
functions because of [C++'s ADL rules](http://en.cppreference.com/w/cpp/language/adl).
55
56
Briefly stated, these are the functions you need to implement, **under
57
your type's namespace**:
58
59
@code{c++}
60
// IMPORTANT: define under the namespace of T. Read note below.
61
namespace your_namespace {
62
63
// tree API implementation for general types (containers
64
// or scalars requiring extra info from the tree):
65
//
66
// needed only if you're deserializing T:
67
c4::yml::ReadResult read(c4::yml::Tree const *tree, c4::yml::id_type node_id, T* var);
68
// needed only if you're serializing T:
69
void write(c4::yml::Tree * tree, c4::yml::id_type node_id, T const& var);
70
71
// or...
72
73
// special case for scalars not needing interaction with the tree:
74
//
75
// needed only if you're deserializing T:
76
bool from_chars(c4::yml::csubstr str, T* var);
77
// needed only if you're serializing T:
78
size_t to_chars(c4::yml::substr buffer, T const& var);
79
// optional:
80
c4::yml::type_bits scalar_flags_val(T const& var); // set extra style flags on T vals
81
c4::yml::type_bits scalar_flags_key(T const& var); // set extra style flags on T keys
82
83
// or...
84
85
// special case for writing string scalars: no need to convert to chars!
86
// mark as string
87
template<> struct c4::is_string<T> : std::true_type {};
88
// instead of to_chars()
89
c4::yml::csubstr to_csubstr(T const& var);
90
// rest as above for scalars
91
92
} // namespace
93
@endcode
94
95
96
@important Because of [C++'s ADL
97
rules](http://en.cppreference.com/w/cpp/language/adl), **it is
98
required to overload these functions in the namespace of the type**
99
you're serializing. Here's an [example of an issue](https://github.com/biojppm/rapidyaml/issues/424)
100
where failing to do this was causing problems in some platforms.
101
102
103
You may also implement %read/write() using the node API instead of the
104
tree API (but read the following section for details):
105
106
@code{c++}
107
// IMPORTANT: define %read() under the namespace of T. Read note above.
108
namespace your_namespace {
109
110
// node API implementation for general types (old approach)
111
// needed only if you're deserializing T:
112
c4::yml::ReadResult read(c4::yml::ConstNodeRef node, T* var);
113
// needed only if you're serializing T:
114
void write(c4::yml::NodeRef &node, T const& var); // can also use NodeRef*, or even NodeRef (by value)
115
116
} // namespace
117
@endcode
118
119
@note For maximum flexibility you should prefer implementing the
120
tree %read/write.
121
122
123
Read on for details.
124
125
126
// <br>
127
// <hr>
128
129
## Why you should prefer implementing with tree API
130
131
You may have noticed above that there are two sets of functions: one
132
for the node API and another for the tree API. You don't need to
133
implement both. Simply put, the choice on which one to implement comes
134
down to which one you want to use, but for maximum flexibility
135
the **default advice is to implement the tree %read/write functions**.
136
137
Here are the key considerations:
138
139
- If you trigger the deserialization from a particular API, it will
140
directly call the corresponding %read/write() function. Further,
141
rapidyaml's default implementation of node is calling into the tree
142
%read/write(), so that if you only implement this one, it is
143
automagically picked even if you're calling from nodes:
144
145
@code{c++}
146
T var;
147
148
// tree calls
149
c4::yml::Tree tree = ...;
150
c4::yml::id_type id = ...; // node id
151
tree.load(id, &var) // calls read(Tree const*,id_type,T*)
152
if(!tree.deserialize(id, &var)) // calls read(Tree const*,id_type,T*)
153
...;
154
tree.save(id, var); // calls write(Tree*,id_type,T const&)
155
tree.set_serialized(id, &var); // calls write(Tree*,id_type,T const&)
156
157
// node calls - forwarding to tree by default
158
c4::yml::NodeRef node = ...;
159
node.load(&var); // calls read(ConstNodeRef const&,T*)
160
// -> rapidyaml calls read(Tree const*,id_type,T*)
161
if(!node.deserialize(&var)) ...; // calls read(ConstNodeRef const&,T*)
162
// -> rapidyaml calls read(Tree const*,id_type,T*)
163
node.save(var); // calls write(NodeRef&,T const&)
164
// -> rapidyaml calls write(Tree*,id_type,T const&)
165
node.set_serialized(&var); // calls write(NodeRef&,T const&)
166
// -> rapidyaml calls write(Tree*,id_type,T const&)
167
@endcode
168
169
- By default, a tree %read/write() impl will get called from a node
170
call. rapidyaml's node impl calls into the tree impl. This means that
171
if you implement the tree %read/write(), rapidyaml will pick it up
172
**even if you are triggering it with the node API**.
173
174
- If you implement node %read/write(), they will be picked up by a
175
node call, but not by a tree call. Further, if you also implement
176
tree %read/writes, they will only be picked up by a tree call.
177
178
- If you implement node %read/write(), it hides rapidyaml's default
179
implementation of calling the tree %read/write(), so if you then
180
want to call tree deserialization, you will also need to implement
181
tree %read/write().
182
183
So again, it is best to choose to implement the tree %read/write() functions.
184
185
186
187
// <br>
188
// <hr>
189
190
## Implementation notes: general types
191
192
As explained above, general types are those that require child nodes
193
(in the case of containers), or are scalars that require extra @ref
194
NodeType flags to be set along with it. For each type, the functions
195
you will to implement depend on whether you're reading or writing from
196
the tree/node.
197
198
199
200
// <br>
201
### Writing general types
202
203
When writing general types to YAML, you need to define the following
204
function:
205
206
@code{c++}
207
// implement these functions for T ...
208
namespace your_namespace { // IMPORTANT read note about namespace above
209
void write(c4::yml::Tree *tree, c4::yml::id_type node_id, T const& var);
210
// or, if you want to use the node API,
211
void write(c4::yml::NodeRef *scalar, T const& var);
212
} // namespace
213
@endcode
214
215
Likewise, for writing keys you need to define the following function
216
(but note the key MUST be a scalar):
217
218
@code{c++}
219
// implement these functions for T ...
220
namespace your_namespace { // IMPORTANT read note about namespace above
221
void write_key(c4::yml::Tree *tree, c4::yml::id_type node_id, T const& var);
222
// or, if you want to use the node API,
223
void write_key(c4::yml::NodeRef *scalar, T const& var);
224
} // namespace
225
@endcode
226
227
The requirements for `%write()` are less numerous than with
228
%read(). Inside `%write()`, you may assume the node is valid, as rapidyaml
229
will have made the required checks before calling your function, as
230
specified by the call triggering the %write (as described in @ref
231
doc_serialization_using).
232
233
As for what you can do inside `%write()`: generally you should only be
234
setting/adding things to the node, and not to its key (that
235
will generally have been dealt with elsewhere), typically with one of
236
[.set_seq()](@ref Tree::set_seq()) or
237
[.set_map()](@ref Tree::set_map()) for containers,
238
or [.set_val()](@ref Tree::set_val()) or
239
[.set_serialized()](@ref Tree::set_serialized()). Following this, for
240
containers you should create and populate the children, with further
241
calls to any of these functions, but now with child nodes and data
242
structures as the targets.
243
244
245
@note See examples of `%write()` implementations:
246
- @ref doc_sample_container_types_brief
247
- @ref doc_sample_container_types
248
- @ref doc_serialization_tree_write
249
- @ref doc_serialization_node_write
250
- see the [vector write implementation](@ref src/c4/yml/std/vector.hpp)
251
- see the [map write implementation](@ref src/c4/yml/std/map.hpp).
252
- see the sample @ref sample_user_container_types
253
- see the sample @ref sample_std_types
254
255
256
257
// <br>
258
### Reading general types
259
260
To enable reading (deserialization) of a custom user type T falling
261
into the general category, you need to define the following function:
262
263
@code{c++}
264
// IMPORTANT: define read() under the namespace of T. Read warning above.
265
namespace your_namespace {
266
c4::yml::ReadResult read(c4::yml::Tree const *tree, c4::yml::id_type node_id, T* var);
267
// and/or, if you prefer the node API
268
c4::yml::ReadResult read(c4::yml::ConstNodeRef node, T* var);
269
} // namespace
270
@endcode
271
272
Likewise, for reading keys you need to define the following function:
273
@code{c++}
274
// IMPORTANT: define %read_key() under the namespace of T. Read warning above.
275
namespace your_namespace {
276
c4::yml::ReadResult read_key(c4::yml::Tree const *tree, c4::yml::id_type node_id, T* var);
277
// and/or, if you prefer the node API
278
c4::yml::ReadResult read_key(c4::yml::ConstNodeRef node, T* var);
279
} // namespace
280
@endcode
281
282
283
Then when you call any of @ref NodeRef::load(), @ref
284
NodeRef::deserialize(), @ref Tree::load() or @ref Tree::deserialize()
285
(as described in @ref doc_serialization_using), rapidyaml will call
286
your `%read()` function through the magic of C++ ADL / Koenig
287
lookup. And likewise, when you call any of @ref NodeRef::load_key(),
288
@ref NodeRef::deserialize_key(), @ref Tree::load_key() or @ref
289
Tree::deserialize_key() (as described in @ref
290
doc_serialization_using), rapidyaml will call your `%read_key()`
291
function. (**But note the rapidyaml tree cannot accept containers as
292
keys!**)
293
294
295
The @ref ReadResult return type is a lightweight truthy type, used to
296
enable reporting either of success or of the offending node, when an
297
error happens in nested reads. It evaluates as true
298
(empty-initialized) when there is no error, or as false on error, and
299
has the innermost node causing the error. This enables accurate error
300
reporting, and is very useful on large YAML files; see also @ref
301
sample_location_tracking() to find the original source location of the
302
offending node.
303
304
305
306
To start with an example, here is the rapidyaml implementation of `%read()` for
307
`std::map`:
308
309
@code{c++}
310
template<class K, class V, class Less, class Alloc>
311
c4::yml::ReadResult read(c4::yml::Tree const* tree, c4::yml::id_type id, std::map<K, V, Less, Alloc> * m)
312
{
313
// RULE 0. you may assume tree and id are valid.
314
if(!tree->is_map(id)) // RULE 1. check node type
315
return c4::yml::ReadResult(id); // report error on this id
316
for(id_type child = tree->first_child(id); child != NONE; child = tree->next_sibling(child))
317
{
318
K k{};
319
// RULE 2. use .deserialize(), not .load()
320
c4::yml::ReadResult result = tree->deserialize_key(child, &k);
321
if((!result))
322
return result; // RULE 3. early exit on error
323
result = tree->deserialize(child, &(*m)[std::move(k)]);
324
if(!result)
325
return result; // may refer to a deeply nested node!
326
}
327
return ReadResult{}; // report success
328
}
329
@endcode
330
331
332
<br>
333
The beginning rule is actually an assumption:
334
335
@important Rule 0. Inside your implementation of `%read()` or
336
`%read_key()`, you may assume the node is valid (ie, that the tree and
337
node_id are valid).
338
339
rapidyaml will already have checked for this as specified by the
340
triggering call (see @ref doc_serialization_using).
341
342
343
<br>
344
Now the first rule:
345
346
@important Rule 1. Inside `%read()`, **start with a node type check**:
347
must be exactly one of @ref VAL (for scalars), @ref SEQ (for sequence
348
types) or @ref MAP (for dictionary types). `%read_key()` *does not
349
require* a @ref KEY check.
350
351
This is needed to ensure that the node type matches the type of the
352
destination variable. Concretely:
353
354
- If you're reading a scalar type like a number or a string, the
355
node must be @ref VAL, ie it must verify @ref NodeType::has_val().
356
357
- If you're reading a sequence type like a vector, the node must be
358
a @ref SEQ, ie it should verify @ref NodeType::is_seq().
359
360
- If you're reading a map type, the node should be a @ref VAL, ie
361
it should verify @ref NodeType::is_map().
362
363
Why can't rapidyaml do this check for you before calling your `%read()`
364
function? Well, in the general case, it is impossible to know what type
365
of node to expect, so rapidyaml can only check that the node is one of
366
the @ref VAL|@ref SEQ|@ref MAP cases above, but not concretely which
367
one. It is up to the `%read()` implementation for a type to specify
368
which one.
369
370
However, note that inside `%read_key()` you do not need a type check,
371
as the rapidyaml tree requires that these are scalars (ie @ref KEY),
372
so rapidyaml does this check for you before calling `%read_key()`.
373
374
375
<br>
376
Now the next rule:
377
378
@important Rule 2. Inside `%read()`, **use
379
[.deserialize()](@ref Tree::deserialize()) and not
380
[.load()](@ref Tree::load())**, to play nice with `.deserialize()`
381
callers calling your function. For `%read_key()` it should be
382
[.deserialize_key()](@ref Tree::deserialize_key()) instead
383
of [.load_key()](@ref Tree::load_key()).
384
385
386
`.load()` triggers an error, while `.deserialize()` just returns, so
387
you don't want to have a `.deserialize()` caller being aborted by a
388
nested `.load()` call in your function. Let the top-level `.load()`
389
caller trigger the error.
390
391
392
<br>
393
Finally,
394
395
@important Rule 3. **Check every read and do early exit on error**,
396
adequately filling the @ref ReadResult return type.
397
398
Your implementation of `%read()` or `%read_key()` **must return a
399
truthy type to signify success of deserialization**. The type should
400
preferably be a @ref ReadResult to enable accurate error reporting.
401
402
If the type is not @ref ReadResult (like the legacy bool), rapidyaml
403
will still work -- although with the inconvenience of pointing only at the
404
outer-most node instead of the actual error-causing node.
405
406
With this return value, rapidyaml will continue on success; on failure
407
it will either return this value to the caller (with `.deserialize()`)
408
or with `.load()` trigger a visit error on the reported node, as
409
instructed by the triggering call (see @ref doc_serialization_using).
410
411
That's it for `%read()`!
412
413
@note See examples of `%read()` implementations:
414
- @ref doc_serialization_tree_read
415
- @ref doc_serialization_node_read
416
- see the [vector read implementation](@ref src/c4/yml/std/vector.hpp)
417
- see the [map read implementation](@ref src/c4/yml/std/map.hpp).
418
- see the sample @ref sample_user_container_types
419
- see the sample @ref sample_std_types
420
421
422
423
424
<br>
425
<hr>
426
427
## Implementation notes: scalars
428
429
When a scalar type does not require any style or tags to be set in the
430
tree, instead of defining `%read()` / `%write()` you can just define
431
the direct serialization functions `%from_chars()` and/or
432
`%to_chars()` to transform the scalar from/to its string
433
representation.
434
435
@note Please take note of the following pitfall when using scalar
436
serialization functions: you may have to include the header with your
437
`%from_chars()` / `%to_chars()` implementation before any other headers
438
that use functions from it.
439
440
441
<br>
442
### Reading scalars
443
444
To implement reading (deserialization) of scalar types, you
445
need to define the following function:
446
447
@code{c++}
448
namespace your_namespace {
449
bool from_chars(c4::yml::csubstr str, T* var); // if you want to read from YAML
450
} // namespace
451
@endcode
452
453
The function receives a string fitted to the scalar, and must convert
454
the string to the argument. To achieve this, you may find it useful to
455
use the utilities in @ref doc_charconv or @ref doc_format, which are
456
very fast and efficient, and play nice with this approach. But that's
457
not mandatory -- you are also free to use any other conversion method
458
you choose, such as fmtlib (but please do not use stringstreams; their
459
performance is really bad).
460
461
Finally, you must return a boolean success status. rapidyaml will then
462
react to this status in accordance with the call triggering the read.
463
464
@note See examples of `%from_chars()` implementations:
465
- for `std::string`: @ref ext/c4core.src/c4/std/string.hpp
466
- for `std::vector<char>`: @ref ext/c4core.src/c4/std/vector.hpp
467
- for `std::span<char>`: @ref ext/c4core.src/c4/std/span.hpp
468
- see the several from_chars overloads in @ref doc_charconv
469
- see the several from_chars overloads in @ref doc_format
470
471
472
<br>
473
### Writing scalars
474
475
To implement writing (serialization) of scalar types, you
476
need to define the following function:
477
478
@code{c++}
479
namespace your_namespace {
480
size_t to_chars(c4::yml::substr buffer, T const& var); // if you want to write to YAML
481
} // namespace
482
@endcode
483
484
This function receives a buffer on which it is to write the
485
serialization of var. Importantly, inside your function **you cannot
486
assume the buffer is large enough** to fit the serialization of
487
var. You must always check against its size.
488
489
You must return the number of bytes required to fit the serialization
490
of var. Importantly, this size must not depend on the size of the
491
buffer, which means **you cannot do an early exit** when you find the
492
buffer is too small. The returned size must be invariant.
493
494
Upon returning, the caller will compare the returned size with the
495
current buffer size. If the returned size is >= than the buffer size,
496
it means the serialization succeeded, and we're done. Otherwise, it
497
means the buffer was too small; then rapidyaml will resize the buffer
498
and call the function again. For an example of this call pattern, see
499
eg @ref serialize_to_arena_scalar().
500
501
A typical implementation of `%to_chars()` will look like this:
502
503
@code{c++}
504
namespace your_namespace {
505
size_t to_chars(c4::yml::substr buffer, T const& var)
506
{
507
size_t pos = 0;
508
for(... var) // iterate over var, adding characters to the buffer
509
{
510
// append another char to the buffer: only if possible!
511
// BUT do not break the loop if the buffer is too small.
512
// Continue doing a blank loop until the end, to count
513
// the needed characters
514
if(pos < buffer.len)
515
buffer[pos] = ...;
516
++pos; // keep counting, even if we already know
517
// the buffer is small!
518
}
519
return pos; // now we know the required size, return it
520
}
521
} // namespace
522
@endcode
523
524
For instance, if your T is a string type, you could do:
525
526
@code{c++}
527
namespace your_namespace {
528
size_t to_chars(c4::yml::substr buffer, T const& var)
529
{
530
size_t sz = var.size();
531
if(sz && sz <= buffer.len)
532
memcpy(buffer.str, var.data(), sz);
533
return sz;
534
}
535
} // namespace
536
@endcode
537
538
@note See examples of `%to_chars()` implementations:
539
- for `std::string`: @ref ext/c4core.src/c4/std/string.hpp
540
- for `std::string_view`: @ref ext/c4core.src/c4/std/string_view.hpp
541
- for `std::vector<char>`: @ref ext/c4core.src/c4/std/vector.hpp
542
- for `std::span<char>`: @ref ext/c4core.src/c4/std/span.hpp
543
- see the several to_chars overloads in @ref doc_charconv
544
- see the several to_chars overloads in @ref doc_format
545
546
547
<br>
548
### Further reading for scalar serialization
549
550
- See the sample @ref sample_user_scalar_types
551
- See the sample @ref sample_formatting for examples
552
of functions from @ref doc_format_utils that will be very
553
helpful in implementing custom @ref to_chars() / @ref from_chars()
554
functions.
555
- See @ref doc_charconv for the example implementations of
556
@ref to_chars() / @ref from_chars() for the fundamental types.
557
- See @ref doc_substr and @ref sample_substr() for the
558
many useful utilities in the substring class.
559
- See quickstart examples on how to @ref doc_sample_scalar_types
560
561
*/
562
563
564
}
// namespace yml
565
}
// namespace c4
c4::yml
Definition
doxy_common.hpp:2
c4
Definition
doxy_common.hpp:1
node.hpp
Node classes.
scalar_charconv.hpp
tree.hpp
doxy_serialization_user_types.hpp
Generated by
1.15.0